Comments

Showing posts sorted by relevance for query gallistel &. Sort by date Show all posts
Showing posts sorted by relevance for query gallistel &. Sort by date Show all posts

Wednesday, January 10, 2018

An update on Empiricism vs Rationalism debates

All papers involve at least two levels of discourse. The first comprises the thesis: what is being argued? What evidence is there for what is being argued? What’s the point of what is being argued? What follows from what is being argued? The second is less a matter of content than of style: how does the paper say what it says? This latter dimension most often reveals the (often more obscure) political dimensions of the discipline, and relates to topics like the following: What parts of one’s butt is it necessary to cover so as to avoid annoying the powers that be? Which powers is it necessary to flatter in order to get a hearing or to avoid the ruinous “high standards of scrutiny” that can always be deployed to block publication? Whose work must get cited and whose can be safely ignored? What are the parameters of the discipline’s Overton Window? If you’ve ever been lucky enough to promote an unfashionable position, you have developed a sensitivity to these “howish” concerns. And if you have not, you should. They are very important. Nothing reveals the core shibboleths (and concomitant power structures) of a discipline more usefully than a forensic inquiry into the nature of the eggshells upon which a paper is treading.  In what follows I would like to do a little eggshell inspection in discussing a pretty good paper from a (for me) unexpected source (a bunch of Bayesians). The authors are Lake, Ullman, Tenenbaum and Gersham (LUTG) all from the BCS group at MIT. The paper (here) is about the next steps forward for an AI that aspires to be cognitively interesting (like the AI of old), and maybe even technologically cutting edge, (though LUTG makes this second point very gingerly).

But that is not all the paper is about. LUTG is also a useful addition to the discussions about Empiricism (E) vs Rationalism (R) in the study of mind, though LUTG does not put matters in quite this way (IMO, for quasi-political reasons, recall “howish” concerns!). To locate the LUTG on this axis will require some work on my part. Here goes.

As I’ve noted before, E and R have both a metaphysical and an epistemological side. Epistemologically, E takes minds to be initially (relatively) unstructured, with mental contour forged over time as a by-product of environmental input as sampled by the senses.  Minds on this view come to represent their environments more or less by copying their structures as manifest in the sensory input.  Successful minds are ones that effectively (faithfully and efficiently) copy the sensory input. As Gallistel & Matzel put it, for Es minds are “recapitulative,” their function being to reliably take “an input that is part of the training input, or similar to it, [to] evoke[s] the trained output, or an output similar to it” (see here for discussion and links). Another way of putting this point is that E minds are pattern matchers, whose job is to track the patternings evident in the input (see here). Pattern matching is recapitulative and E relies on the idea that cognition amounts to tracking the patternings in the sensory input to yield reliable patterns.

Coupled with this E theory of mind comes a metaphysics. Reality has a relatively flat causal structure. There is no rich hierarchy of causal mechanisms whose complex interactions lie behind what we experience. Rather, what you sense is all there is. Therefore, effectively tracking and cataloguing this experience and setting up the right I/O associations suffices to get at a decent representation of what exists. There is no hidden causal iceberg of which this is the sensory/perceptual tip. The I/O relations are all there is. Not surprisingly (this is what philosophers are paid for), the epistemology and the metaphysics fit snugly together.

R combines its epistemology and its metaphysics differently. Epistemologically, R regards sensory innocent minds to be highly structured. The structure is there to allow minds to use sensory/perceptual information to suss out the (largely) hidden causal structures that produce these sensations/perceptions. As Gallistel & Matzel put it R minds are more like information processing systems structured to enable minds to extract information about the unperceived causal structures of the environment that generate the observed patterns of sensation/perception. On the R view, minds sample the available sensory input to construct causal models of the world which generate the perceived sensory patterns. And in order to do this, minds come richly stocked with the cognitive wherewithal necessary to build such models. R epistemology takes it for granted that what one perceives vastly underdetermines what there is and takes it to be impossible to generate models of what there is from sensory perception without a big boost from given (i.e. unlearned) mental structures that make possible the relevant induction to the underlying causal mechanisms/models.

R metaphysics complements this epistemological view by assuming that the world is richly structured causally and that what we sense/perceive is a complex interaction effect of these more basic complexly interacting causal mechanisms. There are hidden powers etc. that lie behind what we have sensory access to and that in no way “resembles” (or “recapitulates”) the observables (see here for discussion).

Why do I rehearse these points yet again? Because I want to highlight a key feature of the E/R dynamic: what distinguishes E from R is not whether pattern matching is a legit mental operation (both views agree that it is (or at least, can be)). What distinguishes them is the Es think that this is all there is (and all there needs to be), while Rs reserve an important place for another kind of mental process, one that builds models of the underlying complex non visible causal systems that generate these patterns. In other words, what distinguishes E from R is the belief that there is more to mental life than tracking the statistical patterns of the inputs. R doesn’t deny that this is operative. R denies that this suffices. There needs to me more, a lot more.[1]

Given this backdrop, it is clear that GG is very much in the R tradition. The basic observation is that human linguistic facility requires knowledge a G (a set of recursive rules). Gs are not perceptually visible (though their products often are). Further, cursory inspection of natural languages indicates that human Gs are quite complex and cannot be acquired solely by induction (even sophisticate discovery procedures that allowed for inductions over inductions over inductions…). Acquiring Gs requires some mental pre-packaging (aka, UG) that enables an LAD to construct a G on the basis of simple haphazard bits of G output (aka, PLD). Once acquired, Gs can be used to execute many different kinds of linguistic behaviors, including parsing and producing novel sentences and phrases. That’s the GG conceit in a small nutshell: human linguistic facility implicates Gs which implicates UGs, both G and UG being systems of knowledge that can be put to multiple uses, including G acquisition, and the production and comprehension of an unbounded number of very different novel linguistic expressions within a given natural language. G, UG, competence vs performance: the hallmarks of a R theory of linguistic minds.

LUTG makes effectively these points but without much discussing the language case (it does at the end but mainly to say that it won’t discuss it).  LUTG’s central critical claim is that human cognition relies on more than pattern matching. There is, in addition, model building, which relies on two kinds of given mental contents to guide learning.  The first kind, what LUTG calls “core” (and an R would call “given”) involves substantive specific content (e.g. naïve physics and psychology). The second involves formal properties (e.g. compositionality and certain kinds of formal analogy (what LUTG calls “learning-to-learn”).[2] LUTG notes that these two features pre-condition further learning and are indispensible. LUTG further notes that the two mental powers (i.e. model building and pattern recognition) can likely be fruitfully combined, though it (rightly, IMO) insists that model building is the central cognitive operation. In other words, LUTG recapitulates the observations above that R theories can incorporate E mechanisms and that E mechanisms alone are insufficient to model human cognition and that without the part left out, E pattern matching models are poor fits with what we have tons of evidence to be central features of cognitive life. Here’s how the abstract puts it:

We review progress in cognitive science suggesting that truly human-like learning and thinking machines will have to reach beyond current engineering trends in both what they learn and how they learn it. Specifically, we argue that these machines should (1) build causal models of the world that support explanation and understanding, rather than merely solving pattern recognition problems; (2) ground learning in intuitive theories of physics and psychology to support and enrich the knowledge that is learned; and (3) harness compositionality and learning-to-learn to rapidly acquire and generalize knowledge to new tasks and situations. We suggest concrete challenges and promising routes toward these goals that can combine the strengths of recent neural network advances with more structured cognitive models.

The paper usefully goes over these points in some detail.  It further notes that this conception of the learning (I prefer the term “acquisition” myself) problem in cognition challenges associationist assumptions which LUTG observes (again rightly IMO) is characteristic of most work in connectionism, machine learning and deep learning. LUTG also points to ways that “model free methods” (aka pattern matching algorithms) might usefully supplement the model building cognitive basics to improve performance and implementation of model based knowledge (see 1.2 and 4.3.2).[3]

Section 3 is the heart of the critique of contemporary AI, which largely ignores the (R) model building ethos that LUTG champions. As it observes, the main fact about humans when compared with current AI models is that humans “learn a lot more from a lot less” (6). How much more? Well, human learning is very flexible and rarely tied to the specifics of the learning situation. LUTG provides a nice discussion of this in the domain of learning written characters. Further, human learning generally requires very little “input.” Often, a couple of minutes or exposures to the relevant stimulus is more than enough. In fact, in contrast to current machine learning or DL systems humans do not need to have the input curated and organized (or, as Geoff Hinton recently put it (here): humans, as opposed to DLers, “clearly don’t need all the labeled data.”). And why not? Because unlike what DLers and connectionists and associationists have forever assumed, human minds (and hence brains) are not E devices but R ones. Or as LUTG puts it (9):

People never start completely from scratch, or even close to “from scratch,” and that is the secret to their success. The challenge of building models of human learning and thinking then becomes: How do we bring to bear rich prior knowledge to learn new tasks and solve new problems so quickly? What form does that prior knowledge take, and how is it constructed, from some combination of inbuilt capacities and previous experience?

In short, the blank tablet assumption endemic to E conceptions of mind is fundamentally misguided and a (the?) central problem of cognitive psychology is to figure out what is given and how this given stuff is used. Amen (and about time)!
Section 4 goes over what LUTG takes to be some core powers of human minds. This includes naïve theories of how the physical world functions and how animate agents operate. In addition, with a nod to Fodor, it outlines the critical role of compositionality in allowing for human cognitive productivity.[4] It is a nice discussion and makes useful comments on causal models and their relation to generative ones (section 4.2.2).
Now let’s note a few eggshells. A central recurring feature of the LUTG discussion is the observation that it is unclear how or whether current DL approaches might integrate these necessary mechanisms. The paper does not come right out and say that DL models will have a hard time with these without radically changing its sub-symbolic associationist pattern matching monomania, but it strongly suggests this. Here’s a taste of this recurring theme (and please note the R overtones).

As in the case of intuitive physics, the success that generic networks will have in capturing intuitive psychological reasoning will depend in part on the representations humans use. Although deep networks have not yet been applied to scenarios involving theory of mind and intuitive psychology, they could probably learn visual cues, heuristics, and summary statistics of a scene that happens to involve agents. If that is all that underlies human psychological reasoning, a data-driven deep learning approach can likely find success in this domain.
However, it seems to us that any full formal account of intuitive psychological reasoning needs to include representations of agency, goals, efficiency, and reciprocal relations. As with objects and forces, it is unclear whether a complete representation of these concepts (agents, goals, etc.) could emerge from deep neural networks trained in a purely predictive capacity. Similar to the intuitive physics domain, it is possible that with a tremendous number of training trajectories in a variety of scenarios, deep learning techniques could approximate the reasoning found in infancy even without learning anything about goal- directed or socially directed behavior more generally. But this is also unlikely to resemble how humans learn, understand, and apply intuitive psychology unless the concepts are genuine. In the same way that altering the setting of a scene or the target of inference in a physics-related task may be difficult to generalize without an understanding of objects, altering the setting of an agent or their goals and beliefs is difficult to reason about without understanding intuitive psychology.

Yup. Right on. But why so tentative? “It is unclear?” Nope it is very clear, and has been for about 300 years since these issues were first extensively discussed. And we all know why. Because concepts like “agent,” “goal,” “cause,” “force,” are not observables and not reducible to observables. So, if they play central roles in our cognitive models then pattern matching algorithms won’t suffice. But as this is all that E DL systems countenance, then DL is not enough. This is the correct point, and LUTG makes it. But note the hesitancy with which it does so. Is it too much to think that this likely reflects the issue mooted at the outset about the implicit politics that one’s dance over the eggshells reveals.
It is not hard to see from how LUTG makes its very reasonable case that it is a bit nervous about DL (the current star of AI). LUTG is rhetorically covering its posterior while (correctly) noting that unreconstructed DL will never make the grade. The same wariness makes it impossible for LUTG to acknowledge its great debt to R predecessors.[5] As LUTG states, its “goal is to build on their [neural networks, NH] successes rather than dwell on their shortcomings” (2). But those that always look forward and never back won’t move forward particularly well either (think Obama and the financial crisis). Understanding that E is deeply inadequate is a prerequisite for moving forward. It is no service to be mealy-mouthed about this. One does not finesse one’s way around veery influential bad ideas.
Ok, so I have a few reservations about how LUTG makes its basic points. That said, this is a very useful paper. It is nice to see this coming out of the very influential Bayesian group at MIT and in a prominent place like B&BS. I am hoping that it indicates that the pendulum is swinging away from E and towards a more reasonable R conception of minds. As I’ve noted the analogies with standard GG practice is hard to miss. In addition LUTG rightly points to the shortcomings with connectionist/deep learning/neural net approaches to mental life. This is good. It may not be news to many of us, but if this signals a return to R conceptions of mind, it is a very positive step in the right direction.[6]


[1] R often goes farther: that even tracking the relevant perceptual regularities requires lots of given mental baggage. The world does not come pre-labeled. So zeroing in on the relevant dimensions for inductive generalization itself requires lots of pre-packaged “knowledge.” Geoff Hinton (the daddy of deep learning) seems to have come to a similar view of late concerning how hard it is to get things off the ground without curated data. See below for a reference.
[2] GGers should find this reminiscent of Chomsky’s discussion in Aspects of substantive and formal universals and how these interact to create Gs (models) of the ambient degenerate and deficient linguistic input available to the child (PLD).
[3] Or to put this in GG friendly terms; LUTG resurrects something very like a competence/performance distinction, model building being analogous to the former and model free methods applied to these models being analogous to the latter. An idea I found interesting is that given models of competence provides a useful domain for performance models wherein sophisticated pattern matching algorithms do their work. Conceptually, this idea seems very similar to what Charles Yang has been advocating as well (see here).
[4] Actually, though compositionality is critical to productivity, it does not suffice for generating an “infinite number of thoughts” or using an “infinite number of sentences”  “from a finite set if primitives” (14). For this we need more than compositionality, we need recursion as well.
[5] It is amazing, IMO, how small a role Chomsky plays in LUTG’s discussion. So far as I can tell, all of its major points were developed in the modern period by him. I am pretty sure that some of the authors know this but that highlighting this fact would hurt the paper politically by offending the relevant leading DL lights.
[6] BTW, LUTG has a nice discussion of the biological “plausibility” of neural net models of the brain. The short point is that short of being pictorially suggestive, there is no real reason for thinking that brains are like connectionist nets. As LUTG puts it (20):

Many seemingly well-accepted ideas concerning neural computation are in fact biologically dubious, or uncertain at best…

For example, most neural networks use some form of gradient based (e.g. back propogation) or Hebbian learning. It has long been argued, however, that
backpropagation is not biologically plausible. As Crick (1989) famously pointed out, backpropagation seems to rquire that information be transmitted backward along the axon, which does not fit with realistic models of neuronal function…

As LUTG observes, this should not in and of itself stop people from investigation neural nets as possible models of brain computation, but it should put an end to the prejudice that brains are nets because they look net-like. Sheesh! 

Saturday, April 7, 2018

Boring details

Bill Idsardi & Eric Raimy



Recall from last time that we are trying to formalize phonology in terms of events (e, f, g, … .; points in abstract time), distinctive features (F, G, … ; properties of events), and precedence, a non-commutative relation of order over events (e^f, etc.). So far, together this forms a directed multigraph.

Inspired by a challenge from Greg Hickok, we’ll see how far we can get with those things without assuming explicit constructs such as sets of features constituting segments (viz subgraphs, more on this in a later post). We will try to avoid falling victim to the pitfalls cautioned against by Kazanina, Bowers & Idsardi 2017 by allowing events to have multiple properties (but we won’t bother collecting up such properties into sets either). Another way to put this point might be to say that we will attempt to cast segments as an emergent phenomena, arising out of features and order (oooh, that sounds so much better). We picked events, features and precedence for inclusion in the model precisely because they are substantive (admittedly, the points in time are pretty abstract), and provide a reasonable starting point for effectual interfaces to action and perception.

Special shit we need:

  • #: an event such that ∀ e NOT e^# (yes, this entails NOT #^#)
  • %: an element such that ∀ e NOT %^e (yes, NOT %^%)
  • $: an empty element, see below. (Think juncture or instantaneous $ilence.)

# codes the beginning of time. (Not the Big Bang, just the abstract time under consideration in the current “workspace”.) % codes the end of time. (And I can’t bring myself to search for a Big Bang Theory quote.) # and % are always available. (I.e. in every workspace instance, see below. We’re building toward MERGE for workspaces. Ideally MERGE(π1,π2) is just the simple union of events, features and relations in both workspaces. But there’s no way it’s quite that simple.)

Another piece of substance that we might want to include is logical relationships among properties. We believe that some of these are universal, but some could also be induced by the learner and so could vary across languages. In particular, for now, we will include the basic feature co-occurrence restrictions and dimensional organization from Avery & Idsardi 1999 (A&I). (In a later post we’ll back off on this and see what being more permissive can buy us.) In the present context this means a couple of things. First, we include in the model the substance of the agonist/antagonist organization of the features (where a NAND b = NOT (a AND b), showing that I [wji] have been disparaging NAND for too long [fn: NAND and NOR are the only logical primitives that alone form a complete basis for the binary logical functions]):

(1) NANDs

  • [spread] NAND [constricted]
  • [stiff] NAND [slack]
  • [raised] NAND [lowered]
  • [high] NAND [low]
  • …

Again, we see this as an empirical claim about phonology. If a good use for e.g. [high, low] events can be found then these conditions should be dropped. (And this is the sort of thing that we will explore later.) These are well-formedness conditions on phonological representations, and would govern rules//laws such that all rules/laws have to obey (1); (1) counts as part of the theory of “possible phonological rule/law”. How they manage to obey (1) is a matter for investigation. 

Second, there are statements like IF F THEN G, meaning  ∀ e IF Fe THEN Ge. (We don’t use the usual symbolic logic symbol for implication, →, because we also want to do graph operations as rules which will also use →.  Another common symbol for implication, ⊃, is also used in set theory, so it has similar problems. So we’ll just spell it out.) Therefore, UG has statements like the following:

(2) IF-THENs

  • IF [spread] THEN [GW]
  • IF [GW] THEN [Laryngeal]
  • ...


If we want to go the whole way to capturing the “bare dimensions” in the A&I framework (and at least one of us does), then we would also have statements like:

(3) XORs

  • IF [GW] THEN [spread] XOR [constricted] 


Given (1), could use OR here instead as F XOR G = (F OR G) AND (F NAND G). (3) would hold at the interface to the motor system but only there, i.e. no bare [GW] sent to the motor system, but bare [GW] is a possible representation in LTM (see A&I) or from the auditory system (= “there was something funky with the voice quality, but I’m not sure what”). This allows for special kinds of underspecification. Anthropomorphizing, a simple [GW] event on its way to the motor interface doesn’t yet know if its [spread] or [constricted] but it will be one or the other (called completion in A&I). 

So far, we believe that this approach is broadly consistent with Jardine 2016. In our terms, Jardine adds another temporal relation (association) between events, we could notate that as e|f, indicating that e is associated to f. (On the “meaning” of association lines see Coleman and Local 1991, Sagey 1986, in the present context we could understand it as weak synchronization, “at about the same time”, recall Saberi & Perrott 1999). In that case there might also be universal restrictions on the combination of these two relations, something like ∀e,f IF e|f THEN NOT e^f AND NOT f^e. But given things like loops in time, this question is not nearly as simple as it first appears. Taking a minimalist position on this question, we will not include | here (but think about ASL again). 

To get theories like Mielke 2008, just start with a different set of features, ones that denote whole “segments,” e.g. [ʊ]e, and then invent/discover/define/select classificatory properties like [round]:

(*4) IF [u] OR [ʊ] OR [o] OR … THEN [round]

Since we’re not going to adopt Mielke’s approach either, we’ve put * in front of 4 to mean that we’re not including it in our model (using the * is actually us trying to trick syntacticians into reading along). Obviously, you could turn such statements around also, IF [round, high, back, ATR, …] THEN [u], but it’s not clear to us what work we would do with [u] if we’ve got the features already. For circuit fanatics, (4) implies a disjunctive relationship between “primary” and “secondary” percepts; if features come “first” then it’s a conjunctive coding in the “second layer” instead. (We're sure the brain does both kinds of things.) 

This brings us to an important point, we have not said anything about how many fneurons (features, properties) an event can have. We could include into the theory statements that make events as simple as possible, meaning that phonological representations would be maximally “scattered”, i.e. have a maximal number of events because each event would have at most one feature. The statement that would do this is:

(*5) ∀ F, G such that F ≠ G, F NAND G 

Statement (5) has an important property, it’s a monadic second order formula because it is  quantifying over properties. This is obscured because we wrote it in a "pointfree" style (the term is from Haskell), so let’s restate it making the event variable clear:

(*6) ∀ e ∀ F, G such that F ≠ G, Fe NAND Ge
(alternatively, ∀ e ∀ F, G IF [F, G]e THEN F = G)

We are quantifying over two different kinds of things, events and monadic properties, and the distinctness criterion (F ≠ G) is over properties, not events (which could be construed extensionally as the set of events Fe and the set of events Ge, but we will resist this interpretation and see properties and functions as first-class, as in Haskell). 

We think that Theory(*6) is an interesting idea to pursue, but we won’t include (*6) here either (we’re fickle like that). That means that events can have multiple properties for us.

Properties in either the articulatory and auditory system can certainly overlap in time. If such “bindings” are recognized by, say, the perceptual system, that information could be conveyed into the phonology as a single event with multiple properties, e.g. [front, round]e. In passing we note that what the phonology receives from perception can be incomplete or distorted in all sorts of ways, but that the phonology trudges on to LTM anyway, despite occlusion (phoneme restoration effects), tinnitus, temporally reversed speech, and so on. 

Consequently, we believe that one aspect of the phonological computation is to construct and deconstruct complex events, that is, to be able to split [front, round]e into [front]e, [round]f, fission, and vice versa (fusion), i.e. between (7) and (8) in either direction:

(7) 


(8) 


To do this we will almost certainly need a derivative relation on events, which is similar to the | we ascribe to Jardine. In a configuration like (8) we would say that e || f (“e is parallel to f”). (This will have to be generalized somewhat, something like a^e AND a^f AND NOT e^f AND NOT f^e.) If we do this generally in a language, then we could possibly capture the effects of line fusion in Government Phonology (Kaye, Lowenstamm and Vergnaud 1985). Another example of something like this might be Korean, in which all fricatives are strident, so we could fuse all parallel strident and fricative events. The net effect of this could be to impair the ability to detect (or possibly even to stably represent) non-strident fricatives. While we’re at this, if a rule can look for an environment like (8) then it’s probably also the case that a learning algorithm could examine such configurations to look for pairs of features that often occur in parallel, to find correlations or to learn event fusion/fission processes.

We return now to $, which we will define as an event without any properties, abstract silence or abstract juncture, a virtual pause. This is mostly for notational convenience, but it could turn out to be used as a basis for some segmentations. For example, sometimes the precedence statements that flow in from the perceptual system will form a bipartite graph between (groups of) events. That is, there will be events a,b,c and x,y,z such one cluster of elements mutually precede the other cluster:

    

The left-hand graph is the graph K3,3, and is one of the two identifiers for non-planar graphs (the other being K5). That is, no graph containing K3,3 can be represented in two dimensions without crossing lines. Perhaps in such cases a $ event e is inserted to establish planarity. (Though why abstract planarity would be important here is not at all clear.) Depending on how common this situation is, this could “chunk” events, or provide alignment possibilities to guide LTM access. 

We could add a clock, and thereby make a timing tier, but we won’t do that either for now. (But we think that such things will ultimately play an important role, see Gallistel & King 2009, and Giraud & Poeppel 2012. The auditory system does have a couple of endogenous clocks, see below.)

If you want to bring SPE back (we are repressing our inner Justin Timberlake here…) then you will need to:

  • add [+F] synonyms for every [F]
  • add all the [-F] definitions
    (IF NOT [+F] THEN [-F] since SPE didn’t allow underspecification), 
  • constrain temporal relations to be linear
    (∀ a, b, e IF a^e AND b^e THEN a=b; IF e^a AND e^b THEN a=b) 
  • constrain temporal relations to be complete
    (∀ e ∃ a, b such that e = # OR e = % OR a^e^b) 

Definitions for things like [αF] are left as an exercise for the reader. Although we like SPE, we won’t do this either.

Also we will include substance to the extent that it will allow for the identification of the articulator-free (manner) features. We’ll assume here that they are privatively coded [stop], [trill], [fricative], [liquid], [approximant], and that these interface to degrees and kinds of innervation to the articulators (i.e. a full ballistic gesture with [stop]). These, along with [nasal] are the landmarks of Stevens 2000 (see also various work by Carol Espy-Wilson and her colleagues), and constitute a proposal for a coarse-coded phonological “primal sketch” (Poeppel & Idsardi 2012). Given that we can have events with multiple properties, we can have events like [stop, Coronal] which would ultimately execute a full closure by the tongue blade. We can also if we want combine the insights of Clements and Rubach on affricates as strident stops with Steriade’s aperture theory. In such cases an event [stop, fric, Coronal, lateral] could be split into two events  [stop, Coronal]^[fric, Coronal, lateral] to interface with motor control. (For those wanting mutually exclusive coding of manner features, i.e. *[stop, fric] = [stop] NAND [fric], just drop the [fric] feature from the “abstract” starting representation for the affricates.) 

Finally, in order to interface with the dual timescale analysis being performed by the auditory system (Poeppel 2003, Giraud & Poeppel 2012), we will allow events which indicate syllables, σe. These events as they come out of the auditory system are innervated along a relatively slower time scale, in theta band (~ 4-8 Hz). The featural elements, like [spread], will be modulated by low gamma band oscillations (~ 20-40 Hz). Together (along with delta band) these form an endogenous clock (see Gallistel and King 2009) which permits the identification of precedence relations inside perception. It is also almost certainly the case that auditory streaming (Bregman 1990) affects the ability to identify precedence relations. Telling that a single neuron displayed an on-off-on pattern (and therefore concluding that ∃ x, y Fx^Fy) is easier than recognizing Fx^Gy for elements across streams, Fink, Ulbrich, Churan & Wittmann 2006. Is this enough by itself to induce “reflective” tier effects in the phonology? Would something like this make tier effects exogenous to the model (the perceptual input to the phonology is just more likely to include x^y statements within a stream)? We don’t think that there are easy answers to these questions.

Given how little specification that we have given to the model, there are few limits on what we can compute with it at the moment. But we are sure that this gives us enough “wiggle room” to capture the observations about linearity, invariance, and biuniqueness in Chomsky 1964.

We want to emphasize here that we do agree with SFP that the computation procedures within phonology are a complex, composed function, which is decomposable into simple functions, though here they are graph-theoretic operations (which are called transductions in that literature) on the events (nodes), properties (features) and precedence relations (links) in the phonological graph. Again, this is like Raimy 2000 without a timing tier, and Autosegmental Phonology without pre-defined tiers or association lines; just let precedence statements be stated between any pair of events, which might have any number of properties.

Next time: Dogs and cats

Friday, May 11, 2018

Ideas that break the mold

I am currently re-reading a terrific book on the history of modern molecular biology called The Eight Day of Creation (here, henceforth 8-day). The book reviews some of the seminal scientific events in modern biology, starting with Watson and Crick’s discovery of the double helix structure for DNA. The book is really fun to read given that it intersperses serious science with lots of titillating gossip about the relevant personalities.

The fun aside, the book (confession: I’ve read the first 200 pages so far and this deals exclusively with DNA) raises two interesting questions for someone like me. 

First, it seems to point to two kinds of “revolutions” in the sciences. The first kind is one that everyone is waiting to happen and that had the work that fomented it not been done, analogous work would soon have been produced making an analogous intellectual contribution. The second kind of work is the opposite: had the people who did it not been around, then nobody else would have done it (or at least not soon). Rather, the idea’s birth would have been long (maybe perpetually) delayed. Both kinds of work are groundbreaking and deserving of the kudos and prizes heaped upon it. The difference is that the discoverers of the first kind are distinguished by breaking the tape a bit ahead of others, while the latter is distinguished by having only one person running the race at all.

The second question, of course, is whether any work currently being pursued in my extended neck of the woods smells like either one of these. What is the next big idea? Needless to say, the first kind will be easier to sniff out than the second given that the second seems to come out of nowhere. But, I suspect that nothing really comes completely out of nowhere and I will suggest that one idea that we have been tracking in FoL that has been treated as scientifically dubious until now is gaining traction so that it is beginning to look like an idea whose time has come. In other words, if we take the progression from Ridiculous! to Obvious! via Sorta/Maybe! as an early indicator of an intellectual revolution, then I think the Gallistel-King conjecture (GKC) is about to enjoy some quality time in the intellectual sun.

It should go without saying (but I will say it nonetheless) that everything I say in what follows is entirely half-assed and speculative. This partially comes with the subject matter. But as I like these sorts of issues, and cannot resist, and have nothing better to talk about at the moment, I will indulge myself. You need not follow.

Let’s start with the two kinds of revolutions. 8-day makes the case (not deliberately, I should add) that the helix was waiting to be discovered and though Watson and Crick got their first, someone else would have grocked the structure very soon if they had not. Likely candidates include Wilkins, Franklin and, almost certainly Pauling. There were probably others around that could have figured out the basic ideas as well (or so 8-day leads me to believe).  In fact, Crick seems to agree with this assessment (see 8-day:155). I do not intend this observation to denigrate the achievement (more exactly: who the hell am I to be able to denigrate it?), just to note that it seems to be an idea whose time had arrived. Many researchers thought that DNA was the important big molecule to chemically understand. They thought this because they knew that it was the repository of hereditary information. Many thought that it was some sort of helix and many thought that X-ray pictures were the right kind of evidence to probe their structure. There were several mathematical accounts available (albeit imperfect) to argue from pictures to structure and it seems from the story 8-day tells that sooner or later the story would be cracked (maybe in dribs and drabs as Crick notes that Medawar suggested in the quoted note on 8-day p. 155).

One could say something similar for other great discoveries. Einstein’s theory of special relativity was very similar to other theories that cropped up at the time (Lorentz, Poincare), Darwin’s theory of natural selection was simultaneously discovered by Wallace. The same appears to be true of the work on QED in more modern times. Again, all of this stuff is great, but it was stuff that seems to have been “in the air” and was something that someone would have discovered pretty soon after whoever is credited with the work did it.[1]

This contrasts with other kinds of discoveries. I am told that Einstein’s General Theory is something that really arrived unexpectedly and that nobody was working along the same lines. Ditto with Mendelian genetics (which was so far ahead of its time that it lay undiscovered for about 35-50 years till it was rediscovered by others (Morgan)). McClintock’s theory of jumping genes might fit in here too from what I know of it as would Marshall and Warren’s theory of the bacterial origins of ulcers (which the rest of the scientific community scoffed at until it received the Nobel). 

To this list, I would add Chomsky’s discovery that humans have an FL built to acquire and use recursive Gs with distinctive computational properties. This is an idea, which though obviously correct is still resisted in many quarters. From my read of the history, it seems clear that had Chomsky not made the case for Generative Grammar nobody would have made it for many years to come (if ever, if current resistance is any indication).

There is one more idea that is coming into its own that I would add to the list, and that brings us to the second question (i.e. anything like this on the horizon now?): Gallistel’s conjecture that human cognitive computation is intra-neuronal and and a species of chemical computation rather than inter-neuronal and “connectionist.” This idea has been roundly resisted (and dismissed) by most of the cog-neuro community. The idea that brain computations are not “like” classical computing at all (no registers, variables, write-to and read-from memory etc.) is a virtual dogma in the neurosciences (and has been for well over 30 years). Neo-connectionism is the name of the cog-neuro game and Gallistel’s critiques have been largely ignored and his more positive proposals barely attended to.  Until recently.

I have noted several recentish studies that have argued that there is (at least) some intra-neuronal calculations that cells do (type “Gallistel-King conjecture” into the find box on the top left corner for posts on the topic). Another one has just appeared in Science(here). The authors  are Tagkopoulos, Liu, and Tavazoie (TLT). They show how the e-coli are capable of “forming internal representations that allow prediction of environmental change.” They do this using “intracellular networks” of “biochemical reactions.” Using these networks, these single cell microbes “form internal representations of their dynamic environments that enable predictive behavior.” Further, consistent with the Gallistel-King conjecture (GKC), it appears that these biochemical representations consist of “genome wide transcriptional responses” based on the DNA-RNA-Protein system characteristic of modern cellular bio-chemistry. 

I am no expert in these matters, but it sure looks like what TLT is finding comports quite nicely with the most straightforward version of the GKC in which cognitive computation is based in the same kinds of processes and networks used to convey hereditary information. First, both take place withinsingle cells. Second the information processing has a pretty classical look and embodies a computational architecture (as discussed in detail in The Gallistel & King book) exploiting DNA/RNA/Proteins in the way GKC initially proposed. Not bad for armchair theorizing. Not bad at all.

As I mentioned, there is more and more stuff coming out that provides empirical support for this big idea. And as I have also mentioned elsewhere, this is roughly what we should expect. The GKC is the conservativehypothesis concerning cognitive computation, despite its also being iconoclastic. It claims that cognition supervenes on an information processing network that we know that cells have and that is used for another purpose (passing traits onto future generations). This system is computationally very rich (it embodies a classical (Turing/von Neumann) computational architecture) as Gallsitel and King show. GKC makes the intellectually conservative proposal that an in placeinformation processing network (aan extant system that passes genetic information across generational time) is also used (or repurposed) for other kinds of info processing tasks (i.e. cognitive information processing). This is standard Darwinian thinking. 

In contrast connectionism is quite radical as it proposes a novel computational apparatus to do the heavy cognitive lifting, bypassing a perfectly respectable extant in place and up and running system. Of course, this might be what happened, but it is still a very radical proposal and should only be accepted if there is very significant evidence in its favor. And as Gallistel has argued, there is really no good evidence to support it and lots of problems with it. I will not rehearse these here (but see here), except to say that it is quite amazing how a bad idea gains staying power if it leverages another really bad idea. The marriage of connectionism and associationism is one such stable couple as Gallistel has shown and the fact the neither is convincing on its own seems not to have convinced the neuro-cognoscenti to dumb the pair.

It is fun to speculate just how game changing a world that accepts GKC would be. Cogneuro could really start stealing liberally from our biological friends. Learning would be to cognition what development is to biology (the building of forms based on genetic information plus environmental inputs). All that inter-neuronal chatter might be re-analyzed as sharing computational results rather than executing actual computations. One could imagine a kind of neuronal wisdom of crowds kind of system where individual neurons compute and then “vote” with the popular favorite output carrying the day. But all of this is realfancifulspeculation, completely unmoored from any knowledge (my specialty!). The important point is that it’s looking more and more like GKC is onto something and if it turns out to be even roughly correct, the consequences for what we do in the cog-neuro sciences will be profound. Why do I think this? Because it’s what happened in biology. Indeed, it’s  the big moral from 8-day. Let me explain.

As I said, 8-day is a terrific read and spurs endless fun speculation. It also carries a moral for linguists (and psychologists) with a cognitive bent. To wit: The intellectual challenge facing people in the mid 50s as regards finding the structure of DNA is quite analogous to the central problems in cog-neuro today. The problem then was to find a way of physically grounding the gene. The problem was usefully bounded by the fact that it had to be a structure that comported with the insights of Mendelian genetics (in particular the fact that reproduction leads to half of the genetic traits of the parents being passed onto the offspring). The intellectual challenge was to find a physical structure that would make clear how this was possible. The Watson-Crick structure for DNA did this in a beautiful way. It showed how Mendel’s genetics could be incarnated. Thus, Mendel’s insights formed a boundary condition on the structure of whatever it was that served to transmit hereditary information. The helical structure of DNA did this almost perfectly and Watson and Crick noted as much in their original paper. Here’s the money quote (8-day:154):

It has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material. 

We find ourselves in a similar situation today. We know a great deal about parts of cognition. We know a lot about some of the computational properties of cognition. We know that this requires representations with complex properties that demand something very like a classical computational architecture. We know that something like innate cognitive knowledge exists that allows for the kinds of cognitive computations biological systems perform. The goal of neuroscience should be to figure out how this is incarnated in biological material: e.g. What’s an address? How do you read-from and write-to the system, what’s a variable? What’s a pointer to an address? How do you store a number in memory? How is “innate” knowledge genetically coded. These are all things we understand how to execute in silicon. The cognitive theories we have tell us that our embodied computational system musthave these kinds of structures and operations as well. The neuro question is how this is embodied in biological material (as opposed to silicon)? The GKC builds on the fact that though we know how this could be done using intra-neuronal chemistry and have no idea how this could be done using inter-neuronal connections. The obvious conclusion is that cognitive computation is physically grounded in intra-neuronal chemistry. Amazingly, we are starting to get some details about how this might done. And nobody would have thunk it when the GKC was first mooted. 


[1]That said, the various “versions” had different virtues. So Einstein’s theory of special relativity was substantially different from Lorentz’s and the ways it was different mattered. So too the various versions of QED. Feynman’s formulation spread quickly to the community because of its intuitive appeal. Schwinger’s (I am told) was no less adequate but it was far more technically challenging and harder to conceptualize. These are not small differences, but the general point stands: the basic analyses were very similar and had Einstein not come up with his theory or Feynman with his someone would have come up with a working version that the community would have embraced.

Tuesday, March 27, 2018

Beyond Epistodome

Note: This post is NOT by Norbert. It's by Bill Idsardi and Eric Raimy. This is the first in a series of posts discussing the Substance Free Phonology (SFP) program, and phonological topics more generally.

Bill Idsardi and Eric Raimy

Before the beginning

For and against method is a fascinating book, documenting the correspondence between Paul Feyerabend and Imre Lakatos in the years just before Lakatos died. Feyerabend proposed a series of exchanges with Lakatos, with Lakatos explicating his Methodology of Scientific Research Programs, and Feyerabend taking the other side, the arguments that became Against Method. We’ll try something similar here, on the Faculty of Language blog, relating to the question of substance in phonology.

Beyond Epistodome

Note: not “Epistemodome” because phonology cares not one whit about etymology Over severalteen posts we will consider the Substance Free Phonology (SFP) program outlined by Charles Reiss and Mark Hale in a number of publications, especially Hale & Reiss 2000, 2008 and Reiss 2016, 2017. We will be concentrating mainly on Reiss 2016 ( http://ling.auf.net/lingbuzz/003087/current.pdf ). Although we agree with many of their proposals, we reject almost all of the rationales they offer for them. Because that’s such an unusual combination of views, we thought that this would be a useful forum for discussion. (And we don’t think any journal would want to publish something like this anyway.) Because this is the Faculty of Language blog (FLog? FoLog? vote in the comments!), we will start with a reading. Today’s reading is from the book of LGB, chapter 1, page 10 (Chomsky 1981):
“In the general case of theory construction, the primitive basis can be selected in any number of ways, so long as the condition of definability is met, perhaps subject to conditions of simplicity of some sort. [fn 12: See Goodman (1951).] But in the case of UG, other considerations enter. The primitive basis must meet a condition of epistemological priority. That is, still assuming the idealization to instantaneous language acquisition, we want the primitives to be concepts that can plausibly be assumed to provide a preliminary, pre-linguistic analysis of a reasonable selection of presented data, that is, to provide the primary linguistic data that are mapped by the language faculty to a grammar; relaxing the idealization to permit transitional stages, similar considerations hold. [fn 13: On this matter, see Chomsky (1975, chapter 3).] It would, for example, be reasonable to suppose that such concepts as “precedes” or “is voiced” enter into the primitive basis …” (emphasis added)
So the motto here is not “substance free”, but rather “substance first, not much of that, and not much of anything else either”. Since we’re writing this during Lent (we gave up sanity for Lent), the message of privation seems appropriate. And we are sure that the minimalist ethos is clear to this blog’s readers as well. Reiss 2016:16-7 makes a different claim:
“• phonology is epistemologically prior to phonetics…
Hammarberg (1976) leads us to see that for a strict empiricist, the somewhat rounded-lipped k of coop and the somewhat spread-lipped k of keep are very different. Given their distinctness, Hammarberg make the point, obvious yet profound, that we linguists have no reason to compare these two segments unless we have a paradigm that provides us with the category k. Our phonological theory is logically prior to our phonetic description of these two segments as “kinds of k”. So our science is rationalist. As Hammarberg also points out, the same reasoning applies to the learner -- only because of a pre-existing built-in system of categories used to parse, can the learner treat the two ‘sounds’ as variants of a category: “phonology is logically and epistemologically prior to phonetics”. Phonology provides equivalence classes for phonetic discussion.” (emphasis added)
Two claims of epistemological priority enter, one claim leaves (or maybe none). The pre-existing built-in system of categories used to parse include: (1) the features (Chomsky, Reiss 2016:18), which they both agree are substantive (Chomsky: “concepts … [that] provide a preliminary pre-linguistic analysis”, Reiss 2016:26: “This work [Hale & Reiss 2003a, 2008, 1998] accepts the existence of innate substantive features”) and (2) precedence (Chomsky; Reiss is mum on this point), also substantive. (We will get to our specific proposal in post # 3.) In the case of the learner, it’s not clear if a claim of epistemological priority can be made in either direction. In our view children have both structures: they come with innate, highly specified motor, perceptual and memory architectures along with a phonology module which has interfaces to those three entities (and probably others besides, as aspects of phonological representations are available for subsequent linguistic processing, Poeppel & Idsardi 2012, and are available in at least limited ways to introspection and metalinguistic judgements, say the central systems of Fodor 1983). The goal for the child is to learn how to transfer information among these systems for the purposes of learning and using the sound structures of the languages that they encounter. We do agree with SFP that a fruitful way of approaching this question is with a system of ordered rules within the phonological component (Halle & Bromberger 1989). In terms of evolutionary (bio-linguistic) priority, it seems blindingly clear that the supporting auditory, motor and memory systems pre-date language, and the phonology module is the new kid on the block. (Whether animal call systems because they also connect memory, action and perception are homologous to phonology is an empirical matter, see Hauser 1996.) In terms of epistemological priority for scientific investigation there would seem to be a couple ways to proceed here (Hornstein & Idsardi 2014). One is to see the human system as primates + X, essentially the evolutionary view, and ask what the minimal X is that we need to add to our last common ancestor to account for modern human abilities. The answer for phonology might be “not much” (Fitch 2018). But there’s another view, more divorced from actual biology, which tries to build things up from first principles. So in this case that would mean asking what can we conclude about any system that needs to connect memory, action, and perception systems of any sort, a “Good Old Fashioned Artificial Intelligence” (GOFAI) approach (Haugeland 1985, see https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence ). This seems to be closer to what Hale & Reiss have in mind, maybe. If so, then by this general MAP definition animal call systems would qualify as phonologies. As would a lot of other activities, including reading-writing, reading-typing, rituals (Staal 1996), dancing, kung-fu fighting, etc. (Nightmares about long ago semiotics classes ensue.) But there are problems (maybe not insurmountable) in proceeding in this way. It’s not clear that there is a general theory of sensation and perception, or of action. And what there is (e.g. Fechner/Weber laws, i.e. sensory systems do logarithms) doesn’t seem particularly helpful in the present context. We think that Gallistel 2007 is particular clear on this point:
“From a computational point of view, the notion of a general purpose learning process (for example, associative learning), makes no more sense than the notion of a general purpose sensing organ—a bump in the middle of the forehead whose function is to sense things. There is no such bump, because picking up information from different kinds of stimuli—light, sound, chemical, mechanical, and so on—requires organs with structures shaped by the specific properties of the stimuli they process. The structure of an eye—including the neural circuitry in the retina and beyond—reflects in exquisite detail the laws of optics and the exigencies of extracting information about the world from reflected light. The same is true for the ear, where the exigencies of extracting information from emitted sounds dictates the many distinctive features of auditory organs. We see with eyes and hear with ears—rather than sensing through a general purpose sense organ--because sensing requires organs with modality-specific structure.” (emphasis added)
So our take on this is that we’re going to restrict the term phonology to humans for now, and so we will need to investigate the human systems for memory, action and perception in terms of their roles in human language, in order to be able to understand the interfaces. But we agree with the strategy of finding a small set of primitives (features and precedence) that we can map across the memory-action-perception (MAP) interfaces and seeing how far we can get with that inside phonology. With Fitch 2018 The phonological continuity hypothesis, though, we will consider the properties of phonology-like systems (especially auditory pattern recognition) in other animals, such as ferrets and finches to be informative about human phonology (see also Yip 2013, Samuels 2015). How much phonological difference does it make that ASL is signed-viewed instead of spoken-heard? Maybe none or maybe a lot, probably some. The idea that there would be action-perception features in both cases seems perfectly fine, though they would obviously be connecting different things (e.g. joint flexion/extension and object-centered angle in signed language and orbicularis oris activation and FM sweep in spoken languages). Does it matter that object-centered properties are further along the cortical visual processing stream (Perirhinal cortex) whereas FM sweeps are identifiable in primary auditory cortex (A1)? Can we ignore the sub-cortical differences between the visual pathway to V1 (simple) and the ascending auditory pathway to A1 (complex)? Does it matter that V1 is two dimensional (retinotopic) and so computations there have access to notions such as spatial frequency that don’t have any clear correlates in the auditory system? Do we need to add spatial relations between features to the precedence relation in our account of ASL? (This seems to be almost certainly yes.) Again, we agree that it’s a good tactic to go as far as we can with features and precedence in both cases, but we won’t be surprised if we end up explanatorily short, especially for ASL. To address a technical point, can you learn equivalence classes? Yes, you can, that’s what unsupervised learning and cluster analysis algorithms do (Hastie, Tibshirani & Freidman 2001). Those techniques aren’t free of assumptions either (No Free Lunch Theorems, Wolpert 1996), but given some reasonable starting assumptions (innate or otherwise) they do seem relevant to human speech category formation (Dillon, Dunbar & Idsardi 2013, see also Chandrasekaran, Koslov & Maddox 2014) even if we ultimately restrict this to feature selection or (de-)activation instead of feature “invention”. Next time: Just my imagination (running away with me)

Thursday, July 17, 2014

Big money, big science and brains

Gary Marcus here discusses a recent brouhaha taking place in the European neuro-science community. The kerfuffle, not surprisingly, is about how to study the brain. In other words, it's about money. The Europeans have decided to spend a lot of Euros (real money!) to try to find out how brains function. Rather than throw lots of it at many different projects haphazardly and see which gain traction, the science bureaucrats in the EU have decided to pick winners (an unlikely strategy for success given how little we know, but bureaucratic hubris really knows no bounds). And, here’s a surprise, many of those left behind are complaining. 

Now, truth be told, in this case my sympathies lie with (at least some) of those cut out.  One of these is Stan Dehaene, who, IMO, is really one of the best cog-neuro people working today.  What makes him good is his understanding that good neuroscience requires good cognitive science (i.e. that trying to figure out how brains do things requires having some specification of what it is that they are doing). It seems that this, unfortunately, is a minority opinion. And this is not good. Marcus explains why.

His op-ed makes several important points concerning the current state of the neuro art in addition to providing links to aforementioned funding battle (I admit it: I can’t help enjoy watching others fighting important “intellectual battles” that revolve around very large amounts of cash). His most important point is that, at this point in time, we really have no bridge between cognitive theories and neuro theories. Or as Marcus puts it:

What we are really looking for is a bridge, some way of connecting two separate scientific languages — those of neuroscience and psychology.

In fact, this is a nice and polite way of putting it. What we are really looking for is some recognition from the hard-core neuro community that their default psychological theories are deeply inadequate. You see, much of the neuro community consists of crude (as if there were another kind) associationists, and the neuro models they pursue reflect this. I have pointed to several critical discussions of this shortcoming in the past by Randy Gallistel and friends (here).  Marcus himself has usefully trashed the standard connectionist psycho models (here). However, they just refuse to die and this has had the effect of diverting attention from the important problem that Marcus points to above; finding that bridge.

Actually, it’s worse than that. I doubt that Marcus’s point of view is widely shared in the neuro community. Why? They think that they already have the required bridge. Gallistel & King (here) review the current state of play: connectionist neural models combine with associationist psychology to provide a unified picture of how brains and minds interact.  The problem is not that neuroscience has no bridge, it’s that it has one and it’s a bridge to nowhere. That’s the real problem. You can’t find what you are not looking for and you won’t look for something if you think you already have it.

And this brings us back to the aforementioned battle in Europe.  Markham and colleagues have a project.  It is described here as attempting to “reverse engineer the mammalian brain by recreating the behavior of billions of neurons in a computer.” The game plan seems to be to mimic the behavior of real brains by building a fully connected brain within the computer. The idea seems to be that once we have this fully connected neural net of billions of “neurons” it will become evident how brains think and perceive. In other words, Markham and colleagues “know” how brains think, it’s just a big neural net.[1] What’s missing is not the basic concepts, but the details. From their point of view the problems is roughly to detail the fine structure of the net (i.e. what’s connected to what). This is a very complex problem for brains are very complicated nets. However, nets they are. And once you buy this, then the problem of understanding the brain becomes, as Science put it (in the July 11/2014 issue), “an information technology” issue.[2]

And that’s where Marcus and Dehaene and Gallistel and a few notable others disagree: they think that we still don’t know the most basic features of how the brain processes information. We don’t know how it stores info in memory, how it retrieves it from memory, how it calls functions, how it binds variables, how, in a word, it computes. And this is a very big thing not to know. It means that we don’t know how brains incarnate even the most basic computational operations.

In the op-ed, Marcus develops an analogy that Gallistel is also fond of pointing to between the state of current neuroscience and biology before Watson and Crick.[3]  Here’s Marcus on the cognition-neuro bridge again:

Such bridges don’t come easily or often, maybe once in a generation, but when they do arrive, they can change everything. An example is the discovery of DNA, which allowed us to understand how genetic information could be represented and replicated in a physical structure. In one stroke, this bridge transformed biology from a mystery — in which the physical basis of life was almost entirely unknown — into a tractable if challenging set of problems, such as sequencing genes, working out the proteins that they encode and discerning the circumstances that govern their distribution in the body.
Neuroscience awaits a similar breakthrough. We know that there must be some lawful relation between assemblies of neurons and the elements of thought, but we are currently at a loss to describe those laws. We don’t know, for example, whether our memories for individual words inhere in individual neurons or in sets of neurons, or in what way sets of neurons might underwrite our memories for words, if in fact they do.

The presence of money (indeed, even the whiff of lucre) has a way of sharpening intellectual disputes. This one is no different. The problem from my point of view is that the wrong ideas appear to be cashing in. Those controlling the resources do not seem (as Marcus puts it) “devoted to spanning the chasm.” I am pretty sure I know why too: they don’t see one. If your psychology is associationist (even if only tacitly so), then the problem is one of detail not principle. The problem is getting the wiring diagram right (it is very complex you know), the problem is getting the right probes to reveal the detailed connections to reveal the full networks. The problem is not fundamental but practical; problems that we can be confident will advance if we throw lots of money at them.

And, as always, things are worse than this. Big money calls forth busy bureaucrats  whose job it is to measure progress, write reports, convene panels to manage the money and the science.  The basic  problem is that fundamental science is impossible to manage due to its inherent unpredictability (as Popper noted long ago). So in place of basic fundamental research, big money begets big science which begets the strategic pursuit of the manageable. This is not always a bad thing.  When questions are crisp and we understand roughly what's going on big science can find us the Higgs field or W bosons. However, when we are awaiting our "breakthrough" the virtues of this kind of research are far more debatable. Why? Because in this process, sadly, the hard fundamental questions can easily get lost for they are too hard (quirky, offbeat, novel) for the system to digest. Even more sadly, this kind of big money science follows a Gresham’s Law sort of logic with Big (heavily monied) Science driving out small bore fundamental research. That’s what Marcus is pointing to, and he is right to be disappointed.




[1] I don’t understand why the failure of the full wiring diagram of the nematode (which we have) to explain nematode behavior has not impressed so many of the leading figures in the field (Cristof Koch is an exception here).  If the problem were just the details of the wiring diagram, then the nematode “cognition” should be an open book, which it is most definitely not. 
[2] And these large scale technology/Big Data projects are a bureaucrats dream. Here there is lots of room to manage the project, set up indices of progress and success and do all the pointless things that bureaucrats love to do. Sadly, this has nothing to do with real science.  Popper noted long ago that the problem with scientific progress is that it is inherently unpredictable. You cannot schedule the arrival of breakthrough ideas.  But this very unpredictability is what makes such research unpalatable to science managers and why it is that they prefer big all encompassing sciency projects to the real thing. 
[3] Gallistel has made an interesting observation about this earlier period in molecular biology. Most of the biochemistry predating Watson and Crick has been thrown away.  The genetics that predates Watson and Crick has largely survived although elaborated.  The analogy in the cognitive neurosciences is that much of what we think of as cutting edge neuroscience might possibly disappear once Marcus’s bridge is built. Cognitive theory, however, will largely remain intact.  So, curiously, if the prior developments in molecular biology are any guide, the cognitive results in areas like linguistics, vision, face recognition etc. will prove to be far more robust when insight finally arrives than the stuff that most neuroscientists are currently invested in.  For a nice discussion of this earlier period in molecular biology read this. It’s a terrific book.