Comments

Monday, September 11, 2017

It can (and has) happened here

So the simple answer seems to be that it both can and does happen in linguistics. The ‘it’ is sexual harassment. An apparently egregious case of such was brought to my attention via a tweet from David Adger directing me to this post. The author is Lauren Hall-Lew (LHL). Her website indicates that she is a sociophonetician working out of Edinburgh.  The post is in reaction to this piece in Mother Jones (MJ), a bit of reporting on the alleged unwanted (and, if true, indefensible) sexual advances of Florian Jaeger on a (then) grad student Celeste King. The MJ piece goes into quite a bit of detail concerning the charges and replies by the University of Rochester (UR) people that looked into this. The most damning bit, IMO, is that the charges were credible enough for Dick Aslin to resign over UR’s tepid response (his assessment). It is also noteworthy that he and Elisa Newport, once chair of the UR cog sci dept (now at Georgeown), were willing to go public seconding Kidd’s claims. These are very serious senior people with impeccable professional reputations and if in their view the stories concerning Jaeger are credible then I take them seriously.  I will not go into the toing and froing that the article reports, but I would like to make a few comments.

First, Kidd is not alone in making the charges. There are others (in fact given the seriousness of the charges, many others) reporting the same thing.

Second, UR’s defense is that they could not prove that what Jaeger did was illegal and would be willing to defend their decision “in a court of law.” This is quite a weak reply. There are many things short of legal culpability that are professionally unacceptable. Moreover, at this moment it seems that UR’s response to the charges has been to harass those faculty that brought the Jaeger behavior to the attention of the higher ups. Here are Aslin and Newport as quoted in MJ:

“The administration has inexplicably failed to defend its most vulnerable citizens—its students—and put future students at risk by failing to act appropriately on their behalf; and it has retaliated against the faculty members whose only motive was to defend these students,” Aslin and Elissa Newport, the former chair of the brain and cognitive sciences department, wrote in a letter delivered to members of the UR Board of Trustees. “The present situation must be viewed as a colossal failure of UR leadership at all levels.”

It is entirely possible that there was no proof sufficient for legal proceedings and yet there was quite a bit of bad behavior.  In fact, the MJ piece suggests that this is indeed what happened. Here is MJ:

The investigation into Jaeger’s behavior took about three months. In her final report, UR investigator Catherine Nearpass concluded that Jaeger had had a sexual relationship with at least one graduate student in the department, as well as a prospective Ph.D. student; that parts of his behavior were inappropriate; and that he “liked to push boundaries with students,” the EEOC complaint alleges. Still, the university ultimately found that Jaeger had not violated the university’s policy against discrimination and harassment, and that there was not enough evidence to conclude he sexually harassed Kidd or any other student in his lab. An appeal was unsuccessful.

The EEOC complaint goes into some detail outlining what Jaeger appears to have done and it ain’t pretty. Might I suggest that it become a policy that lab retreats not include hot tubs and drugs.

Third, it is possible that UR is right, namely that Jaeger’s behavior did not rise to the level of illegality. So what is to be done? This is the topic of LHL’s interesting post. LHL knew Jaeger in grad school and she is not particularly surprised about the charges (adding further evidence that Kidd, Aslin and Newport are onto something). But the real meat of her post, IMO, is a subtle suggestion as to the roots of this behavior:

In discussing this Mother Jones article with other women who went to grad school with us, one made the excellent point that this particular incident has happened because Florian carried his borderline sketchy behaviors from graduate school into postgraduate life, without calibrating for his increase in power relative to the grad students he was interacting with. One grad student flirting with another grad student is one thing. A professor flirting with a grad student is another thing entirely. The options for responding are severely constrained in the latter case in a way they aren’t in the former case. Because power. When someone gains power without checking themselves and their behavior, this is what happens.

I think that this gets it exactly right. The problem isn't merely the behavior but also who is engaged in it. Sexual relations are complicated (as the interlocutor rightly noted). But exploiting a position of power, either implicitly or explicitly, is a bright line.

I would like to push this point a bit further. Jaeger’s alleged bad behavior is rooted in a kind of intellectual dishonesty. The great peculiarly academic vice is seeking deference beyond our domains of competence. Academics are prone to trade on expertise in one area for influence and power in another. What LHL’s interlocutor points out is that when this is coupled with power, this kind of over-reaching goes from often unobjectionable BS to downright disgusting behavior.  Academics of all people should be cognizant when they do this, when they start trading their expertise in one area for generalized power, prestige, kudos, in others. After all, one important intellectual responsibility is knowing when you are speaking authoritatively and when you are full of it. Start confusing the two cases (e.g. start believing you deserve shit because you are smart and know something about one tiny domain of knowledge) and nobody should be surprised that bad behavior follows when given the opportunity that hierarchical relations (student-teacher, mentor-post doc, etc.) offer. So, being blind to power relations is not merely a moral vice (although it is that too), it is also an intellectual failing, and one all too tempting to intellectuals. 

 So, LHL’s colleague has put her finger on an important root cause. Still, what is to be done?  First, kudos again to Newport and (especially) Aslin for putting themselves publically on the line for their students. It is rare to see colleagues publically go on the record regarding the inappropriate behavior of another colleague.  Moreover, their action suggests part of a remedy for the kind of bad behavior Jaeger allegedly engaged in: public shaming! Here’s a suggestion. Stop inviting people who behave like this to speak at your colloquia. Stop inviting them to give special addresses at conferences you are organizing. Stop including them in your workshops. In other words, start shunning them for their bad behavior. 

We don’t generally do this. Why not?  Would we ignore other forms of harassment? Do we really want the people who act this way to become role models for our students? Do we want to expose our students to potential misbehavior?  Isn’t this what we are doing when we pretend that this doesn’t exist?

What the UR events indicated is that we cannot rely on the law to enforce decent behavior on colleagues intent on acting badly.[1] If Aslin’s resignation did nothing, then it is unlikely that this is the most efficient route to reform. But, we can do something. We can make it clear that people who act this way are not going to be honored guests in our departments, conferences, workshops etc. Being invited to these is not someone’s right. Not being invited is not an infringement of academic freedom. It is an honor. It is an indication that the invitee is the kind of person we want to interact with personally and want our students and post-docs to interact with personally. One ingredient in being so honored, is having things worth listening to. But it is only one ingredient. The other is being someone with the right kind of professional integrity and behavior. And in this day and age, this means understanding that acting like Jaeger allegedly did is not acceptable. And one way to make this clear institutionally is not to pretend that it doesn’t exist when we make our decisions about who to honor professionally by inviting them into our professional homes.




[1] Nor am I convinced that legal interventions are always the right solution. The law is a blunt instrument rightly demanding a high level of proof and most usefully applied in the least ambiguous most egregious cases.

Tuesday, September 5, 2017

Explanation, prediction and control

As I never get tired of warning; Empiricism (E) is back and that ain’t good. And it is back in both its guises (see here and here). Let’s review. Oh yes, first an apology. This post is a bit long. It got away from me. Sorry. Ok, some content.

Epistemologically, E endorses Associationist conceptions of mind. The E conception of learning is “recapitulative.” This is Gallistel and Matzel’s (G&M) felicitous phrase (see here). What’s it mean? Recapitulative minds are pattern detection devices (viz.: “An input that is part of the training input, or similar to it, evokes the trained output, or an output similar to it” (see here for additional discussion). G&M contrasts this with (Rish) information processing approaches to learning where minds innately code considerable structure of the relevant learning domain.

Metaphysically, E shies away from the idea that deep casual mechanisms/structures underlie the physical world. The unifying E theme is what you see is what you get: minds have no significant structure beyond what is required to register observations (usually perceptions) and the world has no causal depth that undergirds the observables we have immediate access to.

What unifies these two conceptions is that Eism resists two central Rish ideas (i) that there is significant structure to the mind (and brain) that goes beyond the capacity to record and (ii) that explanation amounts to identifying the hidden simple interacting causal mechanisms that underlie our observations. In other words, for Rs minds are more than experiential sponges (with a wee bit of inductive bias) and the world is replete with abstract (i.e. non observed), simple, casual structures and mechanisms that are the ultimate scaffolding of reality.

Why do I mention this (again!)? It has to do with some observations I made in a previous post (here). There I observed that a recent call to arms (henceforth C&C)[1], despite its apparent self-regard as iconoclastic, was, read one way, quite unremarkable, indeed pedestrian. On this reading the manifesto amounts to the observation that language in the wild is an interaction effect (i.e. the product of myriad interacting factors) and so has to be studied using an “integrated approach” (i.e. by specifying the properties of the interacting parts and adumbrating how these parts interact). This I claimed (and still claim) is a truism. Moreover, it is a universally acknowledged truism, with a long pedigree.

This fact, I suggested, raises a question: Given that what C&C appears to be vigorously defending is a trivial widely acknowledged truism why the passionate defense?[2] My suggestion was that C&C’s real dissatisfaction with GG inclined linguistic investigations had two sources: (i) GGers conviction that grammaticality is a real autonomous factor underlying utterance acceptability and (ii) GGers commitment to an analytic approach to the problems of linguistic performance.

The meaning of (i) is relatively clear. It amounts to a commitment to some version of the autonomy of syntax (grammar) thesis, a weak assumption which denies that syntax is reducible to non-grammatical factors such as parsability, pragmatic suitability, semantic coherence, some probabilistic proxy, or any other non syntactic factor. Put simply, syntax is real and contributes to utterance acceptability. C&C appears to want to deny this weak thesis.

The meaning of (ii) is more fluffy, or, more specifically, it’s denial is. I suggested that C&C seemed attracted to a holistic view that appears to be fashionable in some Machine Learning environs (especially of the Deep Learning (DL) variety). The idea (also endemic to early PDP/connectionist theory) takes the unit of analysis to be the whole computational system and resists the idea that one can understand complex systems by the method of analysis and synthesis. It is only the whole machine/web that computes. This rejects the idea that one can fruitfully analyze a complex interaction effect by breaking it down into interacting parts. Indeed, on this view, there are no parts (and, hence no interacting parts) just one big complex gooey unstructured tamale that does whatever it does.

Here’s a version of this view by Noah Smith (the economist) (3-4):

…deep learning, the technique that's blowing everything else away in a huge array of applications, tends to be the least interpretable of all - the blackest of all black boxes. Deep learning is just so damned deep - to use Efron's term, it just has so many knobs on it. Even compared to other machine learning techniques, it looks like a magic spell…

Deep learning seems like the outer frontier of atheoretical, purely data-based analysis. It might even classify as a new type of scientific revolution - a whole new way for humans to understand and control their world. Deep learning might finally be the realization of the old dream of holistic science or complexity science- a way to step beyond reductionism by abandoning the need to understand what you're predicting and controlling.

This view, is sympathetically discussed here in a Wired article under the title “Alien Knowledge” (AK). In what follows I want to say a little about the picture of science that AK endorses and see what it implies when applied to cog-neuro work. My claim is that whatever attractions it might hold for practically oriented research (I will say what this is anon, but think technology), it is entirely inappropriate when applied to cognitive-neuroscience. More specifically, if one aims for “control” then this picture has its virtues. However, if “understanding” is the goal, then this view has virtually nothing to recommend it. This is especially so in the cog-neuro context. The argument is a simple one: we already have what this view offers if successful, and what it delivers when successful fails to give us what we really want. Put more prosaically, the attraction of the vision relies on confusing two different questions: whether something is the case with how something is the case. Here’s the spiel.

So, first, what's the general idea? AK puts the following forth as the main thesis (1):

The new availability of huge amounts of data, along with the statistical tools to crunch these numbers, offers a whole new way of understanding the world. Correlation supersedes causation, and science can advance even without coherent models, unified theories, or really any mechanistic explanation at all.

There are versions of this claim that are unobjectionable. For many technological ends, we really don’t need to understand mechanism.[3] We once had plenty of evidence that aspirin could work to counteract headaches even if we had no idea how aspirin did this.[4] Ditto, with many other useful artifacts. As Dylan once put it (see here): “you don’t need a weatherman to know which way the wind blows.”  Even less do you need decent meteorological theory. We all know that quite often, for many useful purposes, correlation is more than enough to get you what you want/need. Thus, for many useful practical considerations (even very important ones), though understanding is nice, it is not all that necessary (though see note 4).

Now, Big Data (BD)/Deep Learning (DL) allows for correlations in spades. Big readily available data sets plus DL allows one to find a lot of correlations, some of which are just what the technologist ordered. What makes the above AK quote interesting is that it goes beyond this point of common agreement and proposes that correlation is not merely useful for control and the technology that relies on control, but also for “understanding.” This is newish. We are all familiar with the trope that correlation is not causation. The view here agrees, but proposes that we revamp our conception of understanding by decoupling it from causal mechanism. The suggested revision is that understanding can be had without “models” without “mechanistic explanation,” without “unified theories,” and without any insight into the causal powers that make things tick.  In other words, the proposal here is that we discard the Galilean (and Rish) world view that takes scientific understanding to piggy back on limning the simple underlying causal substructures that complex surface phenomena are manifestations of. Thus, what makes the thesis above interesting is not so much its insistence that correlation (often) suffices for control and that BD/DL together make it possible to correlate very effectively (maybe more so than was ever possible before), but that this “correlation suffices” conception of explanation should replace the old Galilean understanding of the goals of scientific inquiry. Put more baldly, it suggests that the Galilean view is somewhat quaint given the power to correlate that the new BD/DL tools afford.

This vision is clearly driven by a techno imperative. As AK puts it, it’s the standard method of “any self-respecting tech company” and its success in this world challenges the idea that “knowledge [is] about finding the hidden order in chaos [aka: the Galilean world view, NH],” that involves “simplifying the world.” Nope, this is “wrong.” Here’s the money idea: “Knowing the world may require giving up on understanding it” (2).[5]

This idea actually has an old pedigree rooted in two different conceptions of idealization. The first takes it as necessary for finding the real underlying mechanism. This (ultimately Platonic/Rationalist) conception takes what we observe to be a distortion of the real underlying causal powers that are simple and which interact to create the apparent chaos that we see.  The second conception (which is at bottom Eish) takes idealization to distort. For Es, idealizations are the price we pay for our limitations. Were we able to finesse these, then we would not need to idealize and could describe reality directly, without first distorting it. In effect, for Rs, idealization clarifies, for Es it distorts by loosing information.

The two positions have contrasting methodological consequences. For Rs, the mathematization of the world is a step towards explaining it as it is the language best suited for describing and explaining the simple abstract properties/powers/forces that undergird reality.[6] For Es, it is a distortion as there are no simple underlying mechanisms for our models to describe.[7] Models are, at best, compact ways of describing what we see and like all such compactions (think maps, or statistical renderings of lists of facts) they are not as observationally accurate as the thing they compactly describe (see Borges here on maps).

The triumph of modern science is partially the triumph of the Platonic conception as reframed by Galileo and endorsed by Rs. RK challenges this conception, and with it the idea that a necessary step towards understanding is simplification via idealization. Indeed, on the Galilean conception the hardest scientific problem is finding the right idealization. Without it, understanding is impossible. For Platonists/Rs, the facts observed are often (usually) distorted pictures of an underlying reality, a reality that a condign idealization exposes. For Es, in contrast, idealizations distort as through simplifying they set aside the complexity of observed reality.

So why doesn’t idealization distort? Why isn’t simplifying just a mistake, perhaps the best that we have been able to do till now but something to be avoided if possible? The reason is that “good” idealizations don’t distort those features that count: the properties of the underlying mechanisms. What good idealizations abstract away from is the complexity of the observed world. But this is not problematic for the aim is not to “capture the data” (i.e. the observations), but to use some data (usually artificially contrived in artificial controlled settings) to describe and understand the underlying mechanisms that generate the data. And importantly, these mechanisms must be inferred as they are assumed to be abstract (i.e. remote from direct inspection).

A consequence of this is that not all observations are crated equal on the Galilean view (and so capturing the data (whatever that means (which is not at all clear)) is not obviously a scientifically reasonable or even useful project. Indeed, it is part of the view that some data are better at providing a path to the underlying mechanism than other data. That’s why experiments matter; they manufacture useful and relevant data. If this is so, then ignoring some of the data is exactly the way to proceed if one’s aim is to understand the underlying processes. More exactly, if the same forces, mechanisms and powers obtain in the “ideal” case as in the more everyday case and looking at the idealization simplifies the inquiry, then there is nothing wrong (and indeed everything right) with ignoring the more complex case because considering the complex case does not, by hypothesis, shed more light on the fundamentals than does the simpler ideal one.  In other words, to repeat, a good idealization does not distort our perception of the relevant mechanisms though it might only directly/cleanly apply to only a vanishingly small subset of the observable data. The Galilean/R conclusion is that aiming to “cover the data” and “capture the facts” is exactly the wrong thing to be doing when doing serious science.

These are common enough observations in the real sciences. So, for example, Steven Weinberg is happy to make analogous points (see here). He notes that in understanding fundamental features of the physical world everyday phenomena are a bad guide to what is “real.” This implies that successfully attaining broad coverage of commonly available correlations generally tells us nothing of interest concerning the fundamentals. Gravitational attraction works the same way regardless of the shapes of the interacting masses. However, it is much easier to fix the value of the Gravitational constant by looking at what happens when two perfectly spherical masses interact than when two arbitrarily shaped masses do so. The underlying physics is the same in both cases. But an idealization to the spherical case makes life a lot easier.

So too with ideal gasses and ideal speaker-hearers. Thus, in the latter case, abstracting away from various factors (spotty memory, “bad” PLD, non-linguistically uniform PLD, inattention) does not change the fundamental problem of how to project a G from limited “good” “uniform” PLD. So, yes, we humans are not ideal speaker-hearers but considering the LAD to have unbounded memory, perfect attention and learning only from perfect PLD reveals the underlying structure of the computational problem more adequately than does the more realistic actual case. Why? Because the mechanisms relevant to solving the projection problem in the ideal case will not be fundamentally different than what one finds in the actual one, though interactions with other relevant sub-systems will be more complex (and hence more confusing) in the less ideal situation.

So what is the upshot here? Just the standard Galilean/R one: if one aims to understand mechanism, then idealization will be necessary as a way of triangulating on the fundamental principles, which, to repeat, are the object of inquiry.[8]

Of course, understanding the basic principles may leave one wanting in other respects, but this simply indicates that given different aims different methods are appropriate. In fact, it may be quite difficult to de-simplify from the ideal case to make contact with more conventional observations/data. And it is conceivable that tracking surface correlations may make for better control of the observables than is adapting models built on more basic/fundamental/causal principles, as AK observes. Nor should the possibility be surprising. Nor should it lead us to given up the Galilean conception.

But that is not AK’s view. It argues that with the emergence of BD/DL we can now develop systems (programs) that fit he facts without detouring via idealized models. And that we should do so despite an evident cost. What cost? The systems we build will be opaque in a very strong sense. How strong? AK’s premise is that BD/DL will result in systems that are “ineffably complex” (4) (i.e. completely beyond our capacity to understand what causal powers the model is postulating to cover the data). AK stresses that the models that BD/DL deliver are (often) completely uninterpretable, while at the same time being terrifically able to correlate inputs and outputs as desired. AK urges that we ignore the opacity and take this to provide a new kind of “understanding.” AK’s suggestion, then, is that we revise our standards of explanation so that correlation done on a massive enough scale (backed by enough data, and setting up sufficiently robust correlations) just IS understanding and the simplifications needed for Galilean understanding amounts to distortion, rather than enlightenment.  In other words, what we have here is a vision of science as technology: whatever serves to advance the latter suffices to count as an exemplar of the former.

What are the consequences of taking AK’s recommendation seriously? I doubt that it would have a deleterious effect on the real sciences, where the Galilean conception is well ensconced (recall Weinberg’s discussion). My worry is that this kind of methodological nihilism will be taken up within cog-neuro where covering the data is already well fetishized. We have seen a hint of this in an earlier post I did on the cog-neuro of faces reviewing work by Chang and Tsao (see here). Tsao’s evaluation of her work on faces sees it as pushing back against the theoretical “pessimism” within much of neuroscience generated by the caliginous opacity of much of the machine learning-connectionist modeling. Here is Tsao in the NYT:

Dr. Tsao has been working on face cells for 15 years and views her new report, with Dr. Chang, as “the capstone of all these efforts.” She said she hoped her new finding will restore a sense of optimism to neuroscience.

Advances in machine learning have been made by training a computerized mimic of a neural network on a given task. Though the networks are successful, they are also a black box because it is hard to reconstruct how they achieve their result.

“This has given neuroscience a sense of pessimism that the brain is similarly a black box,” she said. “Our paper provides a counterexample. We’re recording from neurons at the highest stage of the visual system and can see that there’s no black box. My bet is that that will be true throughout the brain.”

C&C’s proposals for language study and AK’s metaphysics are part of the same problem that Tsao identifies.  And it clearly has some purchase in cog-neuro or Tsao would not have felt it worth explicitly resisting. So, while physics can likely take care of itself, cog-neuro is susceptible to this Eish nihilism.

Let me go further. IMO, it’s hard to see what a successful version of this atheoretical approach would get us. Let me illustrate using the language case.

Let’s say that I was able to build a perfectly reliable BD/DL system able to mimic a native speaker’s acceptability profile. So, it rated sentences in exactly the way a typical native speaker would. Let’s also assume that the inner workings of this program were completely uninterpretable (i.e. we have no idea how the program does what it does). What would be the scientific value of achieving this? So far as I can tell, nothing. Why so? Because we already have such programs, they are called native speakers. Native speakers are perfect models of, well, native speakers. And there are a lot of them. In fact, they can do much more than provide flawless acceptability profiles. They can often answer questions concerning interpretation, paraphrase, rhyming, felicity, parsing difficulty etc. They really are very good at doing all of this. The scientific problem is not whether this can be done, but how it is getting done. We know that people are very good at (e.g.) using language appropriately in novel settings. The creative aspect of language use is an evident fact (and don’t you forget it! (see here). The puzzle is how native speakers do it, not whether they do. Would writing a program that mimicked native speakers in this way be of any scientific use in explaining how it is done? Not if the program was uninterpretatble, as are those that AK (and IMO, C&C) are advocating.

Of course, such a program might be technologically very useful. It might be the front end of those endless menus we encounter daily when we phone some office up for information (Please listen carefully as our menu has changed (yea right, gotten longer and less useful). Or maybe it would be part of a system that would allow you to talk to your phone as if it were a person and fall hopelessly in love (as in Her). Maybe (though I am conceding this mainly for argument’s sake. I am quite skeptical of all of this techno-utopianism). But if the program really were radically uninterpretable, then its scientific utility would be nil precisely because its explanatory value would be zero.

Let me put this another way. Endorsing the AK vision in, for example, the cog-neuro of language is to confuse two very different questions: Are native speakers with their evident powers possible? versus How do native speakers with their evident powers manage to do what they do? We already know the answer to the first question. Yes, native speakers like us are possible. What we want is an answer to the second. But uninterpretable models don’t answer this question by assumption. Hence the scientific irrelevance of these kinds of models, those that AK and C&C and BD/DLers promote.

But saying this does not mean that the rise of such models might not have deleterious practical effects. Here is what I mean.

There are two ways that AK/C&C/DL models can have a bad effect. The first, as I have no doubt over discussed, is that they lead to conceptual confusions specifically by running together two entirely different questions. The second is that if DL can deliver on correlations the way its hype promises that it will then it will start to dissolve the link between science and technology that has lent popular prestige to basic research. Let me end with a word on this.

In the public mind, it often seems what makes science worthwhile is that it will give us fancy new toys, cure dreaded diseases, keep us forever young and vibrant and clear up our skin. Science is worthwhile because it has technological payoffs. This promise has been institutionalized in our funding practices in the larger impact statements that the NSF and NIH now require as part of any grant application. The message is clear: the work is worthwhile not in itself but to the degree that it cures cancer or solves world hunger (or clears up your skin). Say DL/BD programs could produce perfect correlation machines. These really could be of immense practical value. And given that they are deeply atheoretical, they would drive a large wedge between science and technology. And this would lessen the prestige of the pure sciences because it would lessen their technological relevance. In the limit, this would make theoretical scientific work aiming at understanding analogous to work in the humanities, with the same level of prestige as afforded the output of English and History professors (noooo, not that!!!).

So, the real threat of DL/BD is not that it offers a better answer to the old question of understanding but that it will sever the tie between knowledge and control. This AK gets right. This is already part of the tech start-up ethos (quite different from the old Bell Labs ethos, I might add). And doing this is likely to have an adverse effect on pure science funding (IMO, it has already). Not that I actually believe that technology can propser without fundamental research into basic mechanisms. At least not in the limit case. DL, like the forms of AI that came before, oversells. And like the AI that came before, I believe that it will also crash as a general program, even in technology. The atheoretical (like C&C) like to repeat the Fred Jelinek (FJ) quip that every time he fired a linguist his comp ling models improved. What this story fails to mention is that FJ recanted in later life. Success leveled off and knowing something about linguistics (e.g. sentence and phrasal structure in FJ’s case) really helped out.[9] It seems that you can go far using brute methods, but you then often run into a wall. Knowing what’s fundamentally up really does help the technology along.

One last point (really) and I end. We tend to link the ideas of explanation, (technological) control and prediction. Now, it seems clear that good explanations often license novel predictions and so asking for an explanation’s predictions is a reasonable way of understanding what it amounts to. However, it is also clear that one can have prediction without explanation (that’s what a good correlation will give you even if it explains nothing), and where one has reliable prediction one can also have (a modicum of) control.  In other words, just as AK calls into question the link between understanding and control (more accurately explanation and technology) it also challenges the idea that prediction is explanation, rather than a mark thereof. Explanation requires the right kind of predictions, one’s that follow from an understanding of the fundamental mechanisms. Predictions in and of themselves can be (and in DL models, appear to be) explanatorily idle.

Ok, that is enough. Clearly, this post has got away from me. Let me leave you with a warning: never underestimate the baleful influences of radical Eism, especially when coupled with techno optimism and large computing budgets. It is really deadly. And it is abroad in the land. Beware!



[1] The authors were Christiansen and Chater, ‘C&C’ in what follows refers to the paper, not the authors.
[2] Of course, the authors might have intended this tongue in cheek or thought it a way of staying in shape by exercising in a tilting at windmills sort of way. Maybe. Naaahh!
[3] Note the qualifier ‘many.’ This might be too generous. I suspect that the enthusiasm for these systems lies most extensively in those technological domains where our scientific understanding is weakest.
[4] We now do know what it does (see here). And, interestingly, (as Paul noted in discussion) understanding how it worked allowed us to develop other pain relievers based on different chemical structures (e.g. Ibuprofen). This required getting beyond correlations to understanding mechanism.
[5] The ‘may’ is a weasel word. I assume that we are all grown up enough to ignore this CYA caveat.
[6] And given that mechanism is imperceptible, imagination (the power to conceive of possible unseen mechanisms) becomes an important part of the successful scientific mind.
[7] And hence, for Es imagination is suspect (recall there is no unseen mechanism to divine), the great scientific virtue residing in careful and unbiased observation.
[8] I don’t want to go into this here, but the real problem that Es have with idealization is the belief that there really is no underlying structure to problems. All there is are patterns and all patterns are important and relevant as there is nothing underneath. The aim of inquiry then is to “capture” these patterns and no fact is more relevant than any other as there is nothing but facts that need organizing. The Eish tendency to concentrate on coverage comes honestly, as data coverage is all that there is. So too is the tendency to be expansive regarding what to count as data: everything! If there is no underlying simpler reality responsible for the messiness we see, then the aim of science is to describe the messiness. All of it! This is a big difference with Rs.
[9] See Carl deMarken on the utility of headed phrases for statistical approaches to learning. Btw, Noah Smith (the CS wunderkind at U Wash told me about FJ’s reassessment many years ago)

Friday, September 1, 2017

If this were common knowledge, our funding problems would be over

Every now and then you get a glimpse of another more perfect world. Here is one such. Should Kelly Cherry succeed in convincing the world at large that she is onto something, we would have to beat acolytes off with a stick! And her main focus is sentence diagramming. Imagine the ecstatic joys of long movement, or binding, or even the occasional transgressive island or ECP violation.

So, what starts with an 'S' ends with an 'X' and you do it all night? You got it: SYNTAX. Let's go Kelly!

Monday, August 28, 2017

The normalization of science

The title for this post is meant to suggest Kuhn’s distinction between revolutionary and normal science. The post is prompted by an article in PNAS that Jeff Lidz sent me. It’s by the mathematical brothers Geman. The claim in the opinion note (ON) is that contemporary science is decidedly small bore and lacks the theoretical and explanatory ambitions of earlier scientific inquiry. This despite the fact that there are more scientists doing science and more money spent on research today (I wish that were as obvious in linguistics!) than ever before. ON’s take is that, despite this, today
…advances are mostly incremental, and largely focused on newer and faster ways to gather and store information, communicate, or be entertained. (9384)
Rather than aiming to deliver abstract “unifying theories” concerning basic “mechanisms”, the research challenge is taken to be “more about computation, simulation and “big data”-style empiricism” (9385).
FoLers may recognize that this somewhat jaundiced view fits with my own narrower pet peeves concerning theory within GG. I have complained more than once that theoretical speculation is currently held in low regard. In fact, I believe (and have said so before) that many (in fact, IMO, most) practitioners consider theory to be, at best, useless ornamentation and, at worst, little more than horse doodoo. What is prized is careful description, corralling recalcitrant data points, smoothing well-known generalizations. More general explanatory ambitions are treated with suspicion and held to extremely high standards if given any hearing at all. This, at least, is my view of the current scene (and, IMO, the general hostility towards Minimalist speculation reflects this). The Geman brothers think that the disdain for fundamental unifying theory is part of the larger current scientific ethos. Their note asks what mechanisms drive it.

Before going through their claims, let me point out that neither the Brothers Geman (BG) nor yours truly want to be understood as dissing the less theory driven empirical work that is being done.  Both BG and I appreciate how hard it is to do this work and we also appreciate its importance. That is not the point. Rather, the point is to observe that nowadays only this kind of work is valued and that the field strongly marginalizes theoretical work that has different ambitions (e.g. unification, reduction, conceptual clarification).

OP canvasses several reasons for why this might be so. It considers a few endogenous factors. For example, that the problems scientists tackle today are just harder in that they are “ “unsimplifiable,” not “amenable to abstraction” (9385). OP replies that “many natural phenomena seem mysterious and hopelessly complex before being truly understood.” I would add that the passion for description fits poorly with the readiness to idealize and, if not tempered, it will make the abstraction required for fruitful theorizing impossible. We need to elevate explanatory “oomph” as a virtue alongside data coverage if we are to get beyond ““big data”-style empiricism.”

But, OP does not think that this is the main impetus behind small bore science. It thinks the problems are cultural. This comes in two parts, one of which is pretty standard by now (see here for some discussion and references), and one is more original (or at least I have never considered it). Let me begin with the first more standard observations.

OP believes that scientists today face an incentive system that rewards small bore projects. Fat CVs gain promotion, kudos, grants, and recognition. And fat CVs are best pursued by searching for the “minimal publishable unit” (I loved this term MPUs should become a standard measure) and seeking the best venues for public exposure (viz. wide ranging exposure and publicity being the current coin of the scientific realm). So publish often and be as splashy as possible is what the incentive system encourages and OP thinks that this promotes conservative research strategies that discourage doing something new and different and theoretically novel. Note that this assumes that theoretical novelty is risky in that its rewards are only evident in the longer term. This strikes me as a reasonable bet. However, I think that there is also a tension here: why splashiness encourages conservativity is unclear to me. Perhaps by splashy OP just means getting and remaining in the public eye, rather than doing something truly original and daring.

OP claims that the review process also functions as a conservative mechanism discouraging big ideas:

In academia, the two most important sources of feedback scientists receive about their performance are the written evaluations following the submission of papers for publication and proposals for research funding. Unfortunately, in both cases, the peer review process rarely supports pursuing paths that sharply diverge from the mainstream direction, or even from researchers’ own previously published work. (9386)

As I noted, these two observations are not novel (see here for example), even if they may be well placed. Frankly, I would love to hear from younger colleagues about whether this rings true for them. How deeply do these incentives work to encourage some styles of research and discourage others within linguistics? I think they obtain, but I would love to hear what my younger colleagues think.  I can say from my seat on tenure and promotion committees that CV size matters, though bulk alone is not sufficient. There is a hierarchy of journals and publishing in these is a pre-requisite for hiring and promotion, as are the all important letters. I will say a bit more about this at the end when I comment on OP’s one suggested fix.

OP makes two other cultural observations that I have not seen discussed before concerning how the internet may have changed the conduct of research in unfortunate ways. The first way strikes me as a bid fuddy-duddy in the sense that it sounds like the complaint an old person makes about youngsters.  In fact we hear this claim daily in the popular press regarding the apparent inability of the under 30 to focus given their bad multi-tasking habits. OP carries this complaint over to young researchers who, by constantly being “on-line” and/or “messaging”, end up suffering from a kind of research ADHD. Here is OP (9385):

Less discussed is the possible effect on creativity: Finding organized explanations for the world around us, and solutions for our existential problems, is hard work and requires intense and sustained concentration. Constant external stimulation may inhibit deep thinking. In fact, is it even possible to think creatively while online? Perhaps “thinking out of the box” has become rare because the Internet is itself a box.

This may be true, though I am not sure that I believe it (though being old, I am inclined to believe it). My younger colleagues don’t seem to be that distracted by these new communicative instruments nearly as much as I am. They seem used to it and treat it as just another useful tool. But again, I might be wrong and would love to know from younger colleagues if they think that there is any truth to this.
A second aspect of being connected that OP mentions rings more true to me. Here the issue is not “how we communicate” but “how much” we do so. OP identifies an “epidemic of communication” fed by “easy travel, many more meetings, relentless email and a low threshold for interaction” (9385). OP makes the interesting suggestion that this may be way too much of a good thing. Why so? Because it encourages “cognitive inbreeding.” Here is OP again (9385):

Communication is necessary, but, if there is too much communication, it starts to look like everyone is working in pretty much the same direction. A current example is the mass migration to “deep learning” in machine intelligence.

The first sentence is the useful point. I included the second one for spite because I don’t like Deep Learning and anything that takes a whack at its current fashionability is ok with me. But, the main point is interesting and worth considering. OP even provides a nice analogy with speciation in evolution. Evolution relies on diverse gene pools which requires some isolating of different populations. Too much interaction threatens to homogenize the gene pool and, by analogy, the set of acceptable ideas which in turn makes originality harder. The epidemic of communication also encourages team work, by making collaboration easier. In OP’s opinion, theory is largely a solitary matter and requires an iconoclastic bent of mind, something that is not fostered by an emphasis on team projects and too much collaboration. In place of “big ideas” the new technology fosters “big projects.” That’s the view.

I am not sure that I agree, but it is an intriguing suggestion. It fits with three things I have noticed.
First, that people would rather do anything than think. Thinking is really hard and frustrating. I know that when I am working on something I cannot get my head around I am always looking for something else to read (“can’t get down to the problem until I have mastered all the relevant literature”). I also tidy my office (well, a little). What I don’t do is stare at the problem and stare at the problem and think hard. It’s just too much work, and frustrating. So, I can believe that the ready availability of things to read makes it possible to avoid the hard work of thinking all the more.
Second, and now I am mainly focused on theoretical syntax, we have found no good replacement for Chomsky’s regular revolutions.  Let me be careful here. There is a trope that suggests that GG undergoes a revolution every decade or so and this is used as an indication that GG has made no scientific progress. I think that this is bunk. As I have noted before, I think that our knowledge has accumulated and later theory has largely conserved earlier findings. But, there is a grain of truth to this, but contrary to accepted wisdom the purported “revolutions” have been very good for the field for they have made room for those not invested in the old ideas to advance new ones. In other words, Chomsky’s constantly pulling the rug out from under his old students and shifting attention to new problems, technology and subject matter acted to disrupt complacency and make room for new ways of thinking (to the irritation of the oldsters (and I speak from experience here)). And this was very healthy. In fact, the periods of high theory all started in roughly this way. Chomsky was not the only purveyor of new ideas but he was a reliable source of intellectual disruption. We have, IMO, far less of this now. Rather, we have much more careful filigree descriptive work, but less exciting theoretical novelty. We really need more fights, less consensus. At any rate, this is consonant with OP’s main point, and I find it congenial.

Third, I think that we as a field have come to prize looking busy. In my department for example, it is noticed if someone is not “participating” and we track how many conferences students present at and papers they publish (visible metrics of success that we share with our academic overlords). The idea that a grad student’s main job is to sit and think and play with ideas is considered a bit quaint. Everyone needs to be doing something. But sitting and thinking is not doing in quite the same way that participating in every research group is. It’s harder and more solitary and less valued nowadays, or so it seems. I doubt that this is just true of UMD.

OP makes one more important point, but this sadly is not something we can do much about right now. OP notes that once jobs were very plentiful, as were grants. This made it possible to explore different ideas that might not pan out for your livelihood was not at stake if you swam against the tide or if it too some time before the ideas hit paydirt. I suspect that this is a big cause of the current atmosphere of conformity and timidity that OP identifies. In situations where decisions are made by committees and openings are scarce, the aim is not to offend. Careful, conventional filigree work is safer and playing it safe is a good idea when options are few.

That’s more or less the OP analysis. It also has one suggestion for making things better. It is a small suggestion. Here it is (9386):

Change the criteria for measuring performance. In essence, go back in time. Discard numerical performance metrics, which many believe have negative impacts on scientific inquiry … Suppose, instead, every hiring and promotion decision were mainly based on reviewing a small number of publications chosen by the candidate. The rational reaction would be to spend more time on each project, be less inclined to join large teams in small roles, and spend less time taking professional selfies. Perhaps we can then return to a culture of great ideas and great discoveries.


I like this idea, but I am not sure it will fly. Why? Because it requires that institutions exercise judgment and trust their local colleagues considered opinions. It is easy to count CV entries (it is also “objective” (and hence less liable to abuse)). It’s much harder to evaluate actual research sympathetically and intelligently (and it is also necessarily personal and “subjective” (and so more liable to abuse)). And it is even harder to evaluate evaluations for sympathy and intelligence. At the very least, this is all very labor intensive. I don’t see it happening. What I can see happening is a cut down version of this. We should begin to institutionalize one question (I call it “Poeppel’s question” because he made it the lead off query in everyone of his lab meetings) concerning work we read, review, listen to presentations of, advise: What’s the point?  Why should anyone care? If a demand for an answer to this question becomes institutionalized it will force us all to think more expansively and will promote another less descriptive dimension of evaluation. It’s a small thing, a doable thing. I think it might have surprisingly positive effects. I for one intend to start right away.

Thursday, August 24, 2017

Another sad event

Bill Davies died on August 18. I knew him a bit as he worked some on control and had inquiries concerning some students who were looking for work ay U of I. He was a decent and intelligent man. Here is a more extensive note that Rob Chametzky sent.

*****

It is with profound sadness that we note the passing of Professor William Davies on August 18, 2017. After joining the faculty in 1986, Bill was at the heart of departmental life, serving for many years as Departmental Executive Officer, Director of Graduate Studies, Director of English as a Second Language Programs for thirteen years between 1990 and 2005, and as a mentor and advisor for countless PhD, MA, and BA students in the department. As ESL Director, Bill ensured that ESL Programs’ faculty and staff were treated as respected professionals and as full members of the Department, and that linguistics graduate students pursuing a focus in Teaching English as a Second Language (TESL) were provided teaching experience with structured support and supervision by ESL faculty—just one example, among many, of his strong commitment to student success.
Bill was a highly respected theoretical syntactician and a preeminent scholar of Austronesian languages, focusing extensively on the syntax of raising and control, and on the syntax and morphology of Javanese and Madurese. He received his PhD from the University of California, San Diego in 1981, writing his dissertation on Choctaw clause structure. Beginning with Choctaw, and continuing with other languages, much of Bill's research united his interests and training in syntactic theory with his passion for language documentation and preservation. From the early 1990s, his attraction to the languages of Indonesia drew him first to Javanese, and then to Madurese, a language that he worked on for some twenty years. His work on various linguistic phenomena in Madurese culminated in 2010 in the De Gruyter Mouton A Grammar of Madurese, the first (and only) comprehensive grammar of this language of 14 million speakers. Bill’s theoretical work traversed a broad range of phenomena, including wh-questions, reflexives and reciprocals, antipassives, causatives, and (especially) raising and control. However, his theoretical work was, invariably, coupled with efforts to give back to the people who so graciously allowed him into their space to do his research – studying the grammar of the Madurese while, at the same time, preserving and rendering accessible the rapidly disappearing folk story traditions for their next generation.   
Bill taught and conducted research both at Cornell University, where he held a Mellon postdoctoral fellowship, and at California State University, Sacramento, prior to joining the faculty at Iowa. In addition, he served on the faculty of two Linguistic Society of America Summer Institutes, at the Ohio State University and the University of Chicago, and co-coordinated (with Stan Dubinsky) National Science Foundation-funded workshops at two additional Summer Institutes. Bill made innumerable invited and conference presentations of his research in Indonesia, along with many other international, national, and local venues. He published extensively, both in singly-authored works and in collaborative work with students and colleagues—including, with long-time collaborator Stan Dubinsky, an edited volume on the theory of grammatical functions, two influential volumes on control and raising, and a forthcoming textbook on language conflict and language rights. His research was funded by a number of prestigious grants, including a National Science Foundation grant supporting his work on the grammar of Madurese, and, more recently, grants from the National Endowment for the Humanities, the NSF, and the Smithsonian Institution, as well as the American Institute for Indonesian Studies and the Fulbright Scholar program, for his project documenting the language of the Baduy Dalam. Much of Bill's recent work circled back to his early roots in anthropological linguistics, uniting his interests in culture and language preservation with careful descriptive work, theoretical analyses, and digital audio and video recordings and transcriptions of both Madurese and Baduy folk stories.

During his time on the Iowa faculty, Bill, a gifted and award-winning teacher, particularly enjoyed teaching Linguistic Field Methods and Linguistic Structures— two classes which allowed him to share his passion for Austronesian linguistics with generations of students, and which ultimately spawned many PhD dissertation and qualifying paper topics. To his PhD students, Bill was a tireless mentor and an impeccable role model; he was deeply proud of their accomplishments and cherished this work.

Bill was gentle, warm, compassionate, and funny (and occasionally (okay, often) sarcastic); friends, students, and colleagues will remember his invaluable contributions to the Department of Linguistics and to the University of Iowa, of course, but more so his fundamental decency, his sense of fairness, and his unceasing advocacy for his students and for the field of linguistics. He will be deeply missed.