Comments

Showing posts with label Marr. Show all posts
Showing posts with label Marr. Show all posts

Tuesday, April 5, 2016

Yang on Bayes 1

This is the first of two posts on a recent paper by Charles Yang. Once again, the topic got away from me. I break it down into two parts to prevent you running away too scared to even look. I suspect that my clever maneuver won’t help. But I try.

One more caveat: this is my understanding of Charles’ paper. He should only be held responsible for what I say because he allowed me to read it, and that might be a culpable act.

Charles Yang has a new paper (CY) forthcoming in Language Acquisition and it is a must read. Aside from sketching a very powerful critique of contemporary Bayesianism as applied to linguistic problems (I will return to this critique momentarily), it also does something that I never thought I would witness in my lifetime; it makes quantitative predictions in a linguistic domain. And by “quantitative” I do not mean giving p-values or confidence intervals. I mean numerical predictions about the size of measurable effect. That’s quantitative! So, run, don’t walk to the above link and read the damn thing!

Ok, now that I’ve discharged my kudosing responsibilities, I want to discuss a second feature of CY. It offers an excellent critique of current Bayesian practice as applied to linguistic issues. As you all know, Bayes is a big player nowadays. In fact, I am tempted to say that Bayes has stepped into the position that Connectionism once held as the default theoretical framework in the cognitive sciences in general and the cognition of language in particular. As you may also recall, I have expressed reservations about Bayes and what it brings to the linguistics table (see here, here, here, here). However, it was not until I read CY that I clearly understood what bugs me about the Bayes framework as applied to my little domain of interests. I want to “share” this aha moment with you.

The conclusion can be put briskly as follows: Bayes is not so much wrong as wrong-headed. From a Generative Grammar (GG) perspective, it’s not the details that are off the mark (though they can be) but the idealization implicit in the framework that is (i.e. if you accept as I do that GG has correctly identified the “computational” problems (see here) then Bayes is of little relevance and maybe worse). And because this is so, there is very little we can learn from Bayesian modelings of linguistic problems. There is both (i) not enough there there and (ii) what there there is points in the wrong direction.

Let me put this point another way: all theories idealize. Empirically successful theories built on good idealizations.[1] CY’s argument is that Bayes is a bad idealization when applied to matters linguistic. If it is right (and, surprise surprise, I believe it is) then Bayes not only adds little, it positively misleads and misdirects. Why? Because bad idealizations cannot be empirically redeemed. It’s ok to be wrong (you can build on error given a good framing of the problem). But if a theory is wrongheaded it will impede progress.  Whereas an adequate conception of the problem (which is what idealizations embody) tolerates empirical missteps. You can’t data your way out of a misconceived idealization. Why not? Because if you’ve got the problem wrong then data coverage will come with a big price, ad hoc assumptions whose main purpose is to make up for the basic misconception. Hence, a misframing of the problem leads to explanatory sterility and misplaced efforts. That’s the claim. Let’s see how Bayes, misidealizes when considered form a GG perspective.

A Bayes model consists of 3 moving parts: (i) a hypothesis space delimiting the range of options (possibly weighted), (ii) a specification of the data relevant to the choosing of the right hypothesis, (iii) an update rule saying how to evaluate a given hypothesis given certain evidence. This 3-step procedure can iterate so that data can be evaluated incrementally.[2] What makes a Bayes account distinctive is not this 3-step process. After all, what do (i-iii) say beyond the uncontroversial truism that data is relevant to hypothesis acceptance?  No, what makes the Bayes picture distinctive is how it conceives of the hypothesis space, the relevant data and the update rule. This is what gives Bayes content and this CY argues is where Bayes misfires. It misidentifies the size of the hypothesis space, the amount of relevant data and the nature of the update rule. How exactly? As follows:

a.     Bayes misidentifies the size of the hypothesis space. In particular, Bayes takes as the default assumption that the hypothesis space is large. This makes the basic cognitive problem one of picking out the right hypothesis from a large number of (wrong) possibilities. This is a mis-idealization if within linguistic domains (e.g. acquisition) only a (very) small number of candidates are ever being evaluated wrt the input at any one time.  If the candidate set is small, then the “hard” problem is providing the right characterization of the restricted hypothesis space, not figuring out how to find the right theory in a large space of alternatives.[3]

b.     Bayes mischaracterizes the nature of the data exploited. In particular, Bayes idealizes the learning problem as trying to figure out how to use lots of complex information to find the right hypothesis in a large space. However, CY notes, G acquisition proceeds by using only a small part of the "relevant" data at any given time and there is not much of it. More specifically, the PLD that the child uses is severely restricted (sparse, degenerate and inadequate). Thus, in contrast to the full range of linguistic data that linguists use to find a native speaker’s G, kids only use a severely restricted data set when making zeroing in on their Gs. Thus, for the child, the problem is not finding a G in a large haystack of Gs using a very big pitchfork (that’s the linguist’s problem), but is more like skewering one G from among a small number of scattered Gs using a fragile toothpick (the PLD being multiply deficient). If this is correct, then the hard problem is again finding the structure of the G space so that pretty “weak” evidence (roughly sparsely scattered examples of main clause phenomena)[4] suffices to fix the right G for the LAD.

c.     Bayes uses all the data to evaluate all the hypotheses at every iterative step. The way that Bayes models work is that every hypothesis in the space is evaluated wrt all of the data at any given point (i.e. cross situational learning). So, add new data and we update every hypothesis wrt that data. We might describe this as follows: the procedure takes a panoramic view of the updating function. Contrast this to a procedure where only one theory (or two) is ever seriously being considered at any one step and that alternatives are considered only when the favored one(s) fails (see here). On this second view, evaluation of hypotheses is severely myopic with virtually all but a very few number of alternatives ever being considered. Moreover, the myopia might be yet more severe: the alternatives are not chosen among the most highly valued alternatives but randomly. So, if H1 fails then the procedure does not opt for H2 because it is the next best theory to that point, but the rule is to just pick another hypothesis at random. So, not only is the procedure myopic, it is pretty dumb. Blind and dumb and thus not very rational.

d.     Bayes misunderstands the decision rule to be an optimizing function.  A symptom of this is misunderstanding is the problem Bayes has in accounting for the widespread phenomenon of probability matching (PM). Indeed, Bayes doesn’t explain PM. It explains it away. Why? Because Bayes alone cannot explain it. Bayesian inference is inconsistent with PM. Left to its own devices, Bayes implies that agents will select the option that maximizes the posterior rather than split the difference probabilistically between many different options. But the latter is what PM does (PM implies splitting the difference in accord with the probability of the options). To deal with this (and PM is pervasive), Bayes adds further assumptions (some might argue ad hoc assumptions (see below)) to allow the maximizing rule to result in probability matching. If so, this suggests that the Bayesian idealization without further supplementation points one in the wrong direction (see here for discussion). Much preferred would be an update rule that allows PM. Such rules exist and CY discusses them.

It is worth noting that Bayes makes these substantive assumptions for principled reasons. Bayes’ started out as a normative theory of rationality. Bayes was developed as a formalization of the notion “inference to the best explanation.” In this context the big hypothesis space, full data, full updating, optimizing rule structure noted above are reasonable for they undergird a the following principle of rationality: choose that theory which among all possible theories best matches the full range of possible data. This makes sense as a normative principle of inductive rationality. The only caveat is that Bayes seems to be too demanding for humans. There are considerable computational costs to Bayesian theories (again CY notes these) and it is unclear how good a principle of rationality is if humans cannot apply it. However, whatever its virtues as a normative theory, it is of dubious value in what L.J. Savage (one the founders of modern Bayesianism) termed “small world” situations (SWS).




[1] For example, the view that there is an infinite number of natural language objects is an idealization. It is certainly false that humans can deal manage sentences with, say, 40 levels of embedding. So there is some upper bound on the number of sentences native speakers can effectively manage. Thus we have no behavioral evidence that human linguistic capacity is in fact infinite in the required sense. However, this is not really relevant to the utility of the idealization. What matters is whether it makes the central problem vivid (and tractable) and it does. The basic fact about human linguistic competence is that we can use and understand sentences never before encountered. The question is how we extend beyond what we have been exposed to do this (i.e. the projection problem). It matters not a whit if the domain of our competence extends to an infinite set of objects or to just a very very very large one. Or even small one for that matter. What matters is that we project beyond what we have heard to sentences original to us. If we can do this, we need rules. The infinity assumption vividly highlights the need for rules or Gs in any account of linguistic competence. Thus, in that sense, it is a fine idealization even if, perhaps, false.

I should add, that I am not sure that it is false in the relevant sense. After all, humans have numerical competence even though I doubt that we do well with integers 100,000,000 digits long. My point is that even if false, the idealization is an excellent one for it highlights the problem that needs solving and that whatever solution we come up for making this assumption will extend naturally to the more realistic one. Solving the projection problem in the “larger” domain will serve to solve it in the smaller.
[2] Note, that a procedure is iterative does not mean that it is incremental in the interesting sense that we want from our acquisition models. Incremental in the latter means that the more data we get then the closer we get to the true theory. It means that as we get more data we don’t bounce around the hypothesis space from G to G. Rather as we get more data we smoothly hone in on the right G. It is possible for a theory to be iterative without being incremental. Dresher and Kaye have good discussions of this and the assumption (idealized) that there is instantaneous learning within parameter setting models. The latter idealization makes sense if in fact there is no smooth relation between more data and closing in on the right G (though this is not its only virtue). As we would like our theories to be incremental figuring out how to make this so in, for example, a parameter setting model, is a very interesting question (and the one that Dresher and Kaye focus on).
[3] Moreover, finding the right answer in a restricted space need not be identical to finding the right answer in a large space. Finding your keys on your desk (notice I said “your” not “my”) is not obviously the same problem as finding a needle in a haystack.
[4] CY has an excellent discussion of just how sparse things can be. Standard PoS arguments note how deficient it is relative to much of the knowledge attained.

Saturday, April 2, 2016

Linguistics from a Marrian perspective; an afterword

Big surprise, David Adger and Peter Svenonius got me thinking. Their comments to the two previous posts provoked. Here is a longish attempt to deal with their points. Thx. I urge you to look at their points in detail if you are interested in the Marrian take on these issues.

In standard Marr example, there are non-“mental” magnitudes that “mental” operations are aiming to estimate. So in vision it is shapes, trajectories, parallax, luminescence etc. We have theories of these magnitudes independent of how we estimate them cognitively. They are real physical magnitudes.

So too with addition and prices and cash registers. We understand what addition is and how it works quite independently of how we use it to provide a purchase price.

None of this holds in the language case (or other internal systems in Fodor’s sense).[1] There are no physical magnitudes of relevance or math structures of use. Rather we are trying to zero in on a level 1 theory by seeing how it is used in stylized settings (judgment data being the prime mover here). The judgment task is an interesting probe into the elvel 1 theory because we have reason to believe that it provides a clean picture of the underlying mechanism. Why clean? Because it abstracts away from the exigencies of time pressure, storage pressure, attention pressure etc. that are part and parcel of real time performance. It’s the system functioning at its best because it is functioning well within the limits of its computational (time/space) capacities. That’s the advantage of data drawn in reflective equilibrium. However, this does not mean that it is resource unconstrained (after all any judgment involves parsing, hence memory and attention) but it means that the judgment does not run up against the limits of memory/attention and so likely displays the properties of the computational system more cleanly than if the computational system is cramped by non-computational constraints such as high memory or attention demands.

With this in mind, let’s now get back to a Marr conception of GG.

The creative aspect of language use (that humans can produce and understand an unbounded number of novel sentences) is the BIG FACT that, as Chomsky noted in Current Issues, is one of the phenomena in need of explanation. He notes there that a G is a necessary construct for explaining this obvious behavioral fact (note, that creativity holds is a fact about speakers and what they can do based on what we actually see them do). Without a function that relates sounds and meaning over an unbounded domain (delivers and unbounded number of <s,m> pairs) there is no possible hope of accounting for this behaviorally evident creative capacity. In other words, Gs are necessary for any account of linguistic creativity. 

Here’s a Marr question: what level theory in Marr’s sense is a G in this context? Here’s a proposal: Gs are Marr level 1 theories. If this is correct, we can ask what level 2 theories might look like. Level 2 theories would show how to compute Gish level 1 properties in real time (for processing and production say). So, Gs are level 1 theories and the DTC, for example, is a level 2 theory. The DTC specifies how level 1 constructs are related to  measures of actual on-line sentence processing. If the DTC is correct, it suggests certain kinds of algorithms, one’s that track the derivational complexity of derivations. Of course, none of this is to endorse the DTC (though, to repeat, I do like it quite a bit), but to illustrate Marr-like logic as relates to linguistic theories.

The main problem with Marr's division is not that we can't use it (I just did), but that the explanatory leverage Marr got out of a level 1 theory in vision and cash registers seems absent in syntax. Why? Because Marr’s examples are able to use already available theories for level 1 purposes. In other words, there are already good level 1 theories on offer in vision and cash registers for the finding and these can serve to circumscribe the  the computational problems that must be solved (i.e. explain how the physical optical paramters are mentally computed given input at the retina or how arithmetical functions are embodied in the cash register). Let’s elaborate a bit.

In the vision case, the level 1 account is built on a theory of physical optics which relates objective (non-mental) physical magnitudes (luminescence, shape, parallax, motion, etc.) to info available on the retina. The computational description of the problem becomes how to calculate these real physical magnitudes from retinal inputs. This is a standard inverse problem as there are many ways for these physical variables to relate to patterns of activity on the retina. So the problem is to find the right set of mental constraints in the mental computation that given the retinal input delivers values for the objective variables the level 1 theory specifies. Concepts like “rigidity” serve this purpose to get you shape from motion. Rigidity makes sense as a level 2 computational constraint given the level 1 properties of visual system. So if we assume, for example, objects are rigid then computing their shape using retinal inputs is possible if we can compute their motion from retinal inputs.

In the cash register case in place of optics we have arithmetic. It turns out that calculating grocery bills is a simple math problem with a recognizable arithmetical structure which prices in items embody and that cash registers can calculcate. Given this (i.e. given that we know what the calculation is) we can ask how a cash register does the calculation in real time. How does the cash register do addition? How does it “represent” numbers numerically? Etc.  

None of this is level 1 leverage is available in the language case. Thus, Gs are not constrained by physical magnitudes in the way vision is (the “physics” of language tells us next to nothing about linguistically relevant variables) and there is no interesting math that lies behind syntax (or if there is we haven't found it yet). Linguists need to construct the level 1 theory from scratch and that's what GGers do. The problem does tell us that speakers have internalized recursive procedures (RP) but not the kinds of RPs (and there are endlessly many). It’s the job of GGers to discover the kinds of RPs that native speakers have when they are linguistically able. We argue that our internalized use rules with a certain limited format and generate representations of a certain limited shape. The data we typically use is performance data (judgments) hopefully sanitized to remove many performance impediments (like memory constraints and attention issues). We assume that this data reflects an underlying mental system (or at least I do) that is casually responsible for the judgment data we collect. So we use some cleanish  performance data (i.e. not distorted by sever performance demands) to infer something about the structure of a level 1 theory.

Now if this is the practice, then it looks like it runs together level 1 and level 2 considerations. You cannot judge what you cannot parse. But that's life. We also recognize that delving more deeply into the details of performance might indicate that the level 1 theories we have might need refining (the representations we assume might not be the ones that real time parsing uses, the algorithms might not reflect the derivational complexity of the level 1 theory). Sure. But, and here I am speaking personally, there would be a big payoff if the two lined up pretty closely. Syntactic representations might not be use-representations but it would be surprising to me if the two diverged radically. After all if they did, then how come we pair the meanings we do with the sounds we do? If our stable pairings are due to our G competence then we must be parsing a G function in real time when we judge the way we do. Ditto with the DTC (which I personally believe we have abandoned too quickly, but that’s a story for another time). At any rate, because we don't have (epistemologically) "autonomous" level 1 theories as in vision and cash registers our level 1 and 2 theories are harder to distinguish. Thus, in linguistics, the 1 vs 2 distinction is useful but should not be treated as a dualism. In fact, I take the Pietroski et al work on most to demonstrate the utility of taking the G representation problem to be finding <s,m>s that fit with how we use meanings when actually calculate quantities. How the system engages with other systems during performance can tell us something about the representational format of the system beyond what <s,m> pairings might.

Last point: I can imagine having syntax embodied someplace explicitly or implicitly without being usable. I can even imagine that what we know is in no way implicated in what we do. But I would find this very odd for the linguistic case and even odder given our empirical practice. After all, what we do in practice is infer what we know by looking at what we do in a circumscribed set of doings. This does not imply that we should reduce linguistic knowledge to behavior, but it does seem to imply that our behavior exploits the knowledge we impute and that it is a useful guide to the structure of that knowledge. Once one makes that move, why are some bits of behavior more privileged than others in principle? I can't see why. And if not, then though the competence/performance distinction is useful I would hesitate to confer on it metaphysical substance.

I would actually go a little further: as a regulative ideal we should assume strong transparency between level 1 and level 2 theories in linguistics, though this is not as obvious an assumption to make in the domain of cash registers and vision. I think that it is a very good default assumption that the categories that we think are relevant in our G theories are also the objects our parser parses and our producer produces. There is more to both activities than what Gs describe, but there is at least as much as what Gs describe and in roughly the way that Gs describe it. That’s why judgments are good probes into G structure. So, in our domain, given that we are not in the enviable Marr position of having off the shelf level 1 theories, it is likely that the level 1 theories we develop will be very level 2 pregnant, or so we should assume.

Let me put this another way: say we have two theories that are equally adequate given standard data and say that one (A) fits transparently with our performance theories and the other (B) does not. I would take this as evidence that A is the right level 1 theory. Wouldn’t you? And if you would, then doesn’t this imply that we are taking transparency as a mark of level 1 adequacy? We conclude that the level 1 formats should be responsive to level 2 meshing concerns.

This is not like what we would do in the cash register example (I don’t think). Were we to find that the cash register computes in base 2 rather than base 10 and uses ZF sets as the numerical representation of numbers we would not conclude that it is not “doing” arithmetic. Base 10 or base 2, ZF sets or Arabic numerals it’s doing arithmetic. There is nothing really analogous in the G domain. There might be parsing representations different from G representations, but this is not the default assumption. This makes the Marrian level considerations less clear cut in the language case than the vision case.

To end: thinking Marrishly is a good exercise for the cognitively inclined GGer (hopefully all of you). But, the ling case is not like the others Marr discusses and so we should use his useful distinctions judiciously.




[1] It’s worth recalling Fodor’s thinking that only input systems were modular. Chomsky disagreed. However, what might be right is that only input systems perfectly fit Marr’s 3-level template. This is not surprising given Marr’s interests. As I said in the earlier post, Marr had relatively little to say about higher level object recognition. It is conceivable that there the reason that little progress has been made on this high level topic is the absence of a competence theory in the GG sense.

Wednesday, March 30, 2016

Linguistics from a Marrian perspective 2

This post follows up on this one. The former tries to identify the relevant computational problems that need solving. This one discusses some ways in which the Marrian perspective does not quite fit the linguistic situation. Here goes.

Theories that address computational problems are computational level 1 theories. Generative Grammar (GG) has offered accounts of specific Gs of specific languages, theories of FL/UG that describe the range of options for specific Gs (e.g. GB is such a theory of FL/UG) and accounts that divide the various components of FL into the linguistically specific (e.g. Merge) and the computationally/cognitively general (e.g. Merge vs. feature checking and minimal search). These accounts aim to offer partial accounts for the three questions in (1-3). How do they do this? By describing, circumscribing and analyzing the class of generative procedures that Gs incorporate. If these theories are on the right track, they partially explain how it is that native speakers can understand and produce language never before encountered, what LADs bring to the problem of language acquisition that enables them to converge on Gs of the type they do despite the many splendored poverty of the linguistic input and (this is by far the least developed question) how FL might have arisen from a pre-linguistic ancestor. As these are the three computational problems, these are all computational theories. However, the way linguists do this is somewhat different from what Marr describes in his examples.

Marr’s general procedure is to solve level 1 problems by appropriating some already available off the shelf “theory” that models the problem. So, in his cash register example, he notes that the problem is effectively an arithmetical one (four functions and the integers). In vision the problem is deriving physical values of the distal stimulus given the input stimuli to the visual system. The physical values are circumscribed by our theories of what is physically possible (optics, mechanics) and the problem is to specify these objective values given proximate stimuli. In both cases, well-developed theories (arithmetic, optics) serve to provide ways of addressing the computational problem.

So, for example, the cash register “solves” its problems by finding ways of doing addition, subtraction, multiplication and division of numbers which corresponds to adding items, subtracting discounts, adding many of the same item and providing prices per unit. That’s what the cash register does. It does basic arithmetic. How does it do it? Well that’s the level 2 question. Are prices represented in base 2 or base 10? Are discounts registered on the individual items as registered or is taken off the total at the end? These are level 2 questions of the level 1 arithmetical theory. There is then a level 3 question: how are the level 2 algorithms and representations embodied? Silicon? Gears and fly-wheels? Silly putty and string? But observe that the whole story begins with a level 1 theory that appropriates an off the shelf theory of arithmetic.

The same is true of Marr’s theories of early vision where there are well-developed theories of physical optics to leverage a level 1 theory.

And this is where linguistics is different. We have no off-the shelf accounts adequate to describing the three computational problems noted. We need to develop one and that’s what GG aims to do: specify level 1 computational theories to describe the lay of the linguistic land. And how do we do this? By specifying generative procedures and representations and conditions on operations. These theories circumscribe the domain of the possible; Gs tell us what a possible linguistic object in a specific language is, FL/UG tells us what a possible G is and Minimalist theories tell us what a possible FL is. This leaves the very real question of how the possible relates to the occurrent: how do Gs get used to figure out what this sentence means? How does FL/UG get used to build this G that the LAD is acquiring? How does UG combine with the cognitive and computational capacities of our ancestors to yield this FL (i.e. the ones humans in fact have)? Generative procedures are not algorithms, and (e.g.) representations the parser uses need not be the ones that our level 1 G theories describe.

Why mention this? Because it is easy to confuse procedures with algorithms and representations in Marr’s level 2 sense with Chomsky’s level 1 sense. I know that I confused them, so this is in part a mea culpa and in part a public service. At any rate, the levels must be kept conceptually distinct.

I might add that the reason Marr does not distinguish generative procedures from algorithms or level 1 from level 2 representations is that for him, there is no analogue of generative procedures. The big difference between linguistics and vision is that the latter is an input system in Fodor’s sense, while language is a central system. Early visual perception is more or less pattern recognition, and the information processing problem is to get from environmentally generated patterns to the physical variables that generate these patterns.[1]

There is nothing analogous in language, or at least not large parts of it. As is well known, the syntactic structures we find in Gs are not tied in any particular way with the physical nature of utterances. Moreover, linguistic competence is not related to pattern matching. There are an infinite number of well-formed “patterns,” (a point that Jackendoff rightly made many moons ago). In short, Marr’s story fits input systems better than it does central systems like linguistic knowledge.

That said, I think that the Marr picture presses an issue that linguists should be taking more seriously. The real virtue of Marr’s program for us lies in insisting that the levels should talk to one another. In other words, the work on any level could (and should) inform the theories at the other levels. So, if we know what kinds of algorithms processors use then this should tell us something abut the right kinds of level 1 representations we should postulate.

The work by Pietroski et. al. on most (discussed here) provides a nice illustration of the relevant logic. They argue for a particular level 1 representation of most in virtue of how representations get used to compare quantities in certain visual tasks. The premise is that transparency between level 1 and level 2 representations is a virtue. If it is, then we have an argument that the structure of most looks like this: |{x: D (x) & Y (x)}| > |{ x: D (x)}| - |{x: D(x) & Y (x)}| and not like this: |{x: D (x) & Y (x)}| > {x: D (x) & - Y (x)}|.

Is transparency a reasonable assumption. Sure, in principle. Of course, we may find out that it raises problems (think of the Derivational Theory of Complexity (DTC) in days of yore). But I would argue that this is a good thing. We want our various level theories to inform one another and this means countenancing the likelihood that the various kinds of claims will rub roughly against one another quite frequently. Thus we want to explore ideas like the DTC and representational transparency that link level 1 and level 2 theories.[2]

Let me go further: in other posts I have argued for a version of the Strong Minimalist Thesis (here and here and here) which can be recast in Marr terms as follows: assume that there is a strong transparency between level 1 and level 2 theories in linguistics. Thus, the objects of parsing are the same as those we postulate in our competence theories, and the derivational steps index performance complexity a BOLD responses and other CN measures of occurrent processing and real time acquisition and… This is a very strong thesis for it says that the categories and procedures we discover in our level 1 theories strongly correlate with the algorithms and representations in our level 2 theories. That would be a very strong claim and thus very interesting. In fact, IMO, interesting enough to take as a regulative ideal (as a good research hypothesis to be explored until proven decisively wrong, and maybe even then). This is what Marr’s logic suggests we do, and it is something that many linguists feel inclined to resist. I don’t think we should. We should all be Marrians now.

To end: Marr’s view was that CNers ignored level 1 theories to their detriment. In practice this meant understanding the physical theories that lie behind vision and the physical variables that an information processing account of vision must recover. This perspective had real utility given the vast amount we know about the physical bases of visual stimuli. These can serve to provide a good level 1 theory. There is no analogue in the domain of language. The linguistic properties that we need to specify in order to answer the three computational problems in (1-3) are not tied down in any obvious ways to the physical nature of the “input.” Nor do Gs or FL appear to be all that interesting mathematically so that there is off the shelf stuff that we can use to specify the countours of the linguistic problem. Sure, we know that we need recursive Gs, but there are endlessly many different kinds of recursive systems and what we want for a level 1 linguistic theory is a specification of the one that characterizes our Gs. Noting that Gs are recursive is, scientifically, a very modest observation (indeed, IMO, close to trivial). So, a good deal of the problem in linguistics is that posing the problem does not invite a lot of pre-digested technology that we can throw at it (like arithmetic or optics). Too bad.

However, thinking in level terms is still useful for it serves as a useful reminder that we want our level 1 theories to talk to the other levels. The time for thinking in these terms within linguistics has never been more ripe. Marr 3 level format provides a nice illustration of the utility of such cross talk.


[1] Late vision, the part that gets to object recognition, is another matter. From what I can tell, “higher” vision is not the success story that early vision is. That’s why we keep hearing about how good computers are at finding cats in Youtube videos. One might surmise that the problem vision has with object recognition is that they have not yet developed a good level 1 theory of this process. Maybe they need to develop a notion of a “possible” visual object. Maybe this will need a generative combinatorics. Some have mooted this possibility. See this on “geons.” This kind of theory is recognizably similar to our kinds of GGs. It is not an input theory, though like a standard G it makes contact with input systems when it operates.
[2] Let me once again recommend earlier work by Berwick and Weinberg (e.g. here) that discuss these general issues lucidly.

Monday, March 28, 2016

Linguistics from a Marrian perspective 1

This was intended to be a short post. It got out of hand. So, to make reading easier, I am breaking it into two parts, that I will post this week.

For the cognitively inclined linguist ‘Marr’ is almost as important a name as ‘Chomsky.’ Marr’s famous book (Vision) is justly renowned for providing a three-step program for the hapless investigator. Every problem should be considered from three perspectives: (i) the computational problem posed by the phenomenon at hand, (ii) the representations and algorithms that the system uses to solve the identified computational problem and (iii) the incarnation of these representations and algorithms in brain wetware. Cover these three bases, and you’ve taken a pretty long step in explaining what’s going on in one or another CN domain. The poster child for this Marrian decomposition is auditory localization in the barn owl (see here for discussion and references). A central point of Marr’s book is that too much research eschews step (i), and this has had baleful effects. Why? Because if you have no specification of the relevant computational problem, it is hard to figure out what representations and algorithms would serve to solve that problem and how brains implement them to allow them to do what they do while solving it. Thus, a moral of Marr’s work is that a good description of the computational problem is a critical step in understanding how a neural system operates.[1]

I’m all on board with this Marrian vision (haha!) and what I would like to do in what follows is try to clarify what the computational problems that animate linguistics have been. They are very familiar, but it never hurts to rehearse them. I will also observe one way in which GG does not quite fit into the tripartite division above. Extending Marr to GG requires distinguishing between algorithms and generative procedures, something that Marr with his main interest in early vision did not do. I believe that this is a problem for his schema when applied to linguistic capacities.  At any rate, I will get to that. Let’s start with some basics.

What are the computational problems GG has identified? There are three:

1.     Linguistic Creativity
2.     Plato’s Problem
3.     Darwin’s Problem

The first was well described in the first chapter, first page, second paragraph of Chomsky’s Current Issues. He describes it as “the central fact to which any significant linguistic theory must address.” What is it? The fact that a native speaker “can produce a new sentence of his language on the appropriate occasion, and that other speakers can understand it correctly, though it is equally new to them” (7). As Chomsky goes on to note: “Most of our linguistic experience…is with new sentences…the class of sentences with which we can operate fluently and without difficulty or hesitation is so vast that for all practical purposes (and, obviously, for all theoretical purposes) we may regard it as infinite” (7).

So what’s the first computational problem? To explain the CN sources of this linguistic creativity. What’s the absolute minimum required to explain it? The idea that native speaker linguistic facility rests in part on the internalization of a system of recursive rules that specify the available sound/meaning pairs (<s,m>) over which the native speaker has mastery. We call such rules a grammar (G) and given (1), part of any account of human linguistic capacity must involve the specification of these internalized Gs.

It is also worth noting that providing such Gs is not sufficient. Humans not only have mastery over an infinite domain of <s,m>s, they also can parse them, produce them, and call them forth “on the appropriate occasion.”[2] Gs do not by themselves explain how this gets accomplished, though that there is a generative procedure implicated in all these behaviors is as certain as anything can be once one recognizes the first computational problem.

The second problem, (2), shifts attention from the properties of specific Gs to how any G get acquired. We know that Gs are very intricate objects. They contain some kinds of rules and representations and not others. Many of their governing principles are not manifest in simple data of the kind that it is reasonable to suppose that children have easy access to and that they can easily use. This means that Gs are acquired under conditions where the input is poor relative to the capacity attained. How poor? Well, the the input is sparse in many places, degraded in some, and non-existent in others.[3]  Charles Yang’s recent Zipfian observations (here) demonstrate how sparse the input is even in seemingly simple cases like adjective placement. Nor is the input optimal (e.g. see how sub-optimal word “learning” is in real world contexts (here and here)). And last, but by no means least, for many properties of Gs there is virtually zero relevant data in the input to fix their properties (think islands, ECP effects, and structure dependence).

So what’s the upshot given the second computational problem? G acquisition must rely on given properties of the acquirer that are instrumental to the process of G acquisition. In other words, the Language Acquisition Device (LAD) (aka, child) comes to the task of language acquisition with lots of innate knowledge that the LAD crucially exploits in acquiring its particular G. Call this system of knowledge the Faculty of Language (FL). Again, that LADs have FLs is a necessary part of any account of G acquisition. Of course, it cannot be the whole story and Yang (here) and Lidz (here) (a.o.) have offered models of what more might be involved. But, given the poverty of the linguistic stimulus relative to the properties of the G attained, any adequate solution to the computational problem (2) will be waist deep in innate mental mechanisms.

This leaves the third problem. This is the “newest” on the GG docket, and rightly so, for its investigation relies on (at least partial) answers to the first two. The problem addressed is how much of what the learner brings to G acquisition is linguistically specific and how much is cognitively and/or computationally general. This question can be cast in computational terms as follows: assume a pre-linguistic primate with all of the cognitive and computational capacities this entails, what must be added to these cognitive/computational resources to derive the properties of FL? Call the linguistically value added parts “Universal Grammar” (UG). The third question comes down to trying to figure out the fine structure of FL; how much of FL is UG and how much generic computational and cognitive operations?

A little thought places interesting restrictions on any solution to this problem. There are two relevant facts, the second being more solid than the first.

The first one is that FL has emerged relatively recently in the species (sourly 100kya) and when it emerged it did so rapidly. The evidence for this is “fancy culture” (FC). Evidence for FC consists of elaborate artifacts/tools, involved rituals, urban centers, farming, forms of government etc. and these are hard to come by before about 50kya (see here). If we take FC as evidence for linguistic facility of the kind we have, then it appears that FL emerges on the scene within roughly the last 100k years.

The second fact is much more solid. It is clear that humans of diverse ethnic and biological lineage have effectively the same FL. How do we know? Put a Piraha in Oslo and it will develop a Norwegian G at the same rate and trajectory as other Norwegians do and with the same basic properties. Ditto with a Norwegian in the forests of the Amazon living with the Piraha. If FL is what underlies G acquisition, then all people have the same basic FL given that anyone of them could acquire any G if appropriately situated. Or, whatever FL is, it has not changed over (at least) the last 50ky. This makes sense if the emergence of FL rested on very few moving parts (i.e. it was a “simple” change).[4] 

Given these boundary conditions, the solution to the Darwin’s problem must bottom out on an FL with a pretty slight UG; most of the computational apparatus of FL being computationally and cognitively generic.[5] 

So three different computational problems, which circumscribe the class of potential solutions. How’s this all related to Marr? And this is what is somewhat unclear, at lest to me. I will explain what I mean in the next post.



[1] The direction of inference is not always from level 1 to 2 then to 3. Practically, knowing something about level 2 could inform our understanding of the level 1 problem. Ditto wrt level 3. The point is that there are 3 different kinds of questions one can use to decompose the CN problem, and that whereas level 2 and 3 questions are standard, level 1 analyses are often ignored to the detriment of the inquiry. But I return to the issue of cross-talk between levels at the end.
[2] This last bit, using them when appropriate is somewhat of a mystery. Language use is not stimulus bound. In Chomsky’s words, it is “appropriate to circumstance without being caused by them.” Just how this happens is entirely opaque, a mystery rather than a problem in Chomsky terminology. For a recent discussion of this point (among others) see his Sophia lectures in Sophia Linguistica #64 (2015).
[3] Charles Yang’s recent work demonstrates how sparse it is even in seemingly simple cases like adjective placement.
[4] It makes sense if what we have now is not the result of piecemeal evolutionary tinkering for if it were the result of such a gradual process it raises the obvious question of why the progress stopped about 50kya. Why didn’t FLs further develop to advantage Piraha to acquire Piraha and Romance speakers to acquire Romance? Why stop with an all purpose FL when one more specialized to the likely kind of language the LAD would be exposed to was at hand? One answer is that this more bespoke FL was never on offer; all you get is the FL based on the “simple” addition or nothing at all. Well, we all got the same one.
[5] So, much of the innate knowledge required to acquire Gs from PLD is not domain specific. However, I personally doubt that there is nothing proprietary to language. Why? Because nothing does language like we do it, and given its obvious advantages, it would be odd if other animals had the wherewithal to do it but didn’t. Sort of like a bird that could fly never doing so. Thus, IMO, there is something special about us and I suspect that it was quite specific to language. But, this is an empirical question, ultimately.