Comments

Showing posts with label efp. Show all posts
Showing posts with label efp. Show all posts

Monday, April 9, 2018

Arbitrary, or dogs and cats

Bill Idsardi

In recent comments Veno has put one part of the SFP view very succinctly and clearly:


"In other words, phonology (as an aspect of the mind/brain) treats features (and other units of phonological representation) as arbitrary symbols. From the point of view of phonology, then, features are substance-free units. This, of course, does not mean that features are not related to phonetic substance, and such a conceptualization of features does not preclude the construction of a neurobiologically plausible interface theory (even spelled out in Marr’s terms)."

So I think we need to unpack what "arbitrary symbols" means here. To put my cards on the table, when I read "arbitrary" I think of two things: the use of arbitrary in mathematics (meaning "anything meeting the definition") and Saussure's arbitrariness of the sign. My worry is that this view tends to exclude an important middle, the existence of abstraction -- distinct, but non-arbitrary relationships between levels of representation, call it hidden substance (because it's useful and partly veridical). And I think such non-arbitrary relationships of abstraction play an important role in sensory systems, and constitute the "substance" (or "substantive relationship") between adjacent levels of representation (being veridical and useful). So what we're going to uncover in this blog post are a few instances of "hidden substance". A word of warning -- this post might not make for light bedtime reading. On the other hand, it might be very soporific.

Let's start with a simple example of a non-arbitrary system, the unary numeral system, or tally marks. In this system for the natural numbers, 0 is the empty string, 1 = "|", 2 = "||", 3 = "|||" and so on. This is a non-arbitrary system because "more is more": larger numbers are represented with larger representations (larger numbers correspond to larger data structures). And addition in this system is concatenation, which then automatically (and non-arbitrarily) preserves important properties like being associative and commutative. So this isn't purely Saussurean as the representations have hidden substance (or partial substance if you prefer). Onomatopoeia (sound-symbolism) is another kind of in-between case, and it might be helpful at least as an analog in understanding the point here, namely that there's still some substance (veridicality and usefulness) but it has been partly obscured by the mapping (but this analogy is imperfect, and I really don't want to discuss theories of sound-symbolism here). It is also obscured by the fact that many other mappings are much more arbitrary (e.g. Roman numerals). Perhaps another way to think about this idea would be to say that arbitariness can be put on a scale of how much of the system is done via lookup tables, and how much of the system is done by general laws of combination (see Gallistel and King 1999).

Sensory systems, even mechano- and chemo-transducers, tend to do a similar kind of abstraction in lawfully transmitting some relationships. But admittedly those cases are not nearly as clean as the unary numerals toy example. For example, the conversion done by the rod cells in the retina also abides by "more is more" -- within the operational limits more photons received means more activity. (At the low end the limit reaches down to a single photon, at the high end, the rods reach saturation pretty quickly, leaving the rods relatively less to do for humans in the modern built world.)

I think the intended use of arbitrary in Veno's quote is for something like "substitutable, interchangeable in the functions, operations and/or relations". This is also what I take to be the point of the dogs-cats argument, which is in Bridget's article that she mentioned in response to our first post. Here's Bridget quoting Daniel Currie Hall, channeling Alec Marantz.


"The phonological component does not need to know whether the features it is manipulating refer to gestures or to sounds, just as the syntactic component does not need to know whether the words it is manipulating refer to dogs or to cats; it only needs to know that the features define segments and classes of segments. The phonetic component does not need to be told whether the features refer to gestures or to sounds, because it is itself the mechanism by which the features are converted into both gestures and sounds. So it does not matter whether a feature at the interface is called [peripheral], [grave], or [low F2], because the phonological component cannot differentiate among these alternatives, and the phonetic component will realize any one of them as all three." (p 206)

I have a couple of comments about this argument. First, I think that  the appropriate comparison for phonology is semantics, not syntax, because it is semantics that connects to the CI interface. That is, the question is whether the difference between dog() and cat() is semantically substantive, not if they are syntactically distinct. But the argument is presented in various places ranging across both semantics and syntax. And to be clear, I think this observation about cats and dogs is correct about both syntax and semantics. That is, interchanging cat() for dog() doesn't affect the nature of the syntactic or semantic computations, though it might eventually end up in different truth values for particular instances (e.g. dog(laika) vs cat(laika)). (And this is not an endorsement on my part of truth-value-oriented semantics, see Pietroski, but it will do for this discussion.)

The issue I have with the argument is simply that the range of semantic examples (cats vs dogs) is too narrow to reveal the hidden semantic substance. These cases both have the same type, (e,t). The important question is not about dogs and cats but about dogs and some item with a different type, like all which has type ((e,t), ((e,t),t)). Is the difference between dogs and all important within the semantic computation? The answer would seem to be yes -- that's the whole point of having a type system. And consequently, the difference between dogs and all is again a kind of hidden substance, only now between the CI notions DOG and ALL that map to dogs and all. That is, the difference between DOG and ALL in the CI system maps to a type difference in the semantics between (e,t) and ((e,t), ((e,t),t)). Consequently we should restrict the notion of free substitution to substitution between items of the same type. For me, this interchangeability notion is not a question of being "substance-free" (although one might say that the calculations are substance-agnostic, adapting Thomas Graf's use of "content-agnostic" from his reply to Norbert's post), rather the idea seems more closely related to the notion of systematicity (Fodor and Pylyshyn 1988, Aizawa 2003).

So let's try to translate the free-substitution-within-types idea into the proposed EFP structure to try to find cases of hidden substance. The resulting claim, which we largely agree with, would say that all events are freely interchangeable with other events without affecting the nature of the computation, all features are interchangeable with other features, and all precedence relations are interchangeable with other precedence relations. Not, importantly, that there are no differences between events and features and precedence relations, which all have different type signatures (in a "truth-value" phonology they would be (e), (e,t) and (e,(e,t)) respectively). So, we'll say a little boldly, type differences indicate hidden substantive differences. But having hidden substance isn't being "substance free" though it does make it harder to spot.

Free substitutability is certainly a situation "devoutly to be wished", even when restricted to items of the same type, but I'm afraid we will still fall a little short of this ideal. Why? Because there are the dreadful "special cases", meaning more examples of hidden substance. In the EFP model, the # and % events are special cases, which means that they aren't fully interchangeable with other (ordinary) events. An automatic consequence of their special status is that the precedence relations involving # and % are not fully interchangeable with other precedence relations either. And furthermore that there are no features that apply to # and % events. (I.e. [spread]e for e = # isn't a thing, though see this exchange between Lass and Morris Halle.) One way of thinking about this is that # and % are "purely formal", but that again highlights that the difference is one of substance -- either the special events don't have any (which seems not quite right), or they have weird, special properties that are preserved across the interface ("beginning/end of form").

This is similar in some respects to math cases where certain Abelian groups are extended to form fields. In those cases we need to "special case" the identity element over addition (0) to say that it does not have a multiplicative inverse (i.e. 0 has no reciprocal, or you can't divide by 0). I think in general we are so used to these special cases that we often fail to even notice them. So let me point out a very general source of special cases. In a recursive definition, the special cases will be the base cases (= stopping cases), like "end of string" or 0. Are there any additional special cases for events, features or precedence beyond the ones just noted? Unfortunately, we suspect that there are some more. Again, the minimalist program qua program is to try to keep the special cases to a minimum, not to declare them all out-of-bounds a priori.

To bring this post to another bumper-sticker conclusion, the hidden substance cases show that arbitrary ≠ systematic ≠ substance-free. We need to keep these notions separate, because they interact in interesting ways in complex, modular systems like language.

Wednesday, April 4, 2018

A modest proposal

Bill Idsardi and Eric Raimy


[Note: the following owes tremendous debts to Jeff Heinz and Paul Pietroski, but that does NOT imply their endorsement. But they can provide endorsements or disavowals in the comments if they want to. They also have the right to remain silent, because ‘Murica.]

Taking phonology to be the mind/brain model for speech, it needs to have interfaces to at least three other systems: the articulatory motor control system (action), the auditory system (perception) and the long term memory system (memory), forming a Memory-Action-Perception (MAP) loop (Poeppel & Idsardi 2012). Likewise, signed languages will have to have interfaces to action (the motor system for the hands, arms, etc.), perception (the visual system) and to memory. As discussed in the comments last week, we think therefore that sign language phonology probably also includes spatial primitives which spoken language phonology lacks. 

So the data structures inside phonology proper must be able to effectively receive and send information across those interfaces. This condition is a basic tenet of the minimalist program. Pietroski (2003: 198 Chomsky and his critics) is on-point and unmistakable :

"Indeed, a sentence is said to be a pair of instructions -- called "PF" and LF" -- for the A/P and C/I systems. If these instructions are to be usable by the extralinguistic systems, PFs and LFs may have to respect constraints that would be arbitrary from a purely linguistic perspective."

(And see Chomsky's replies in the same book for further remarks on the importance of the interface conditions, e.g. p 275.) That is, the data structures inside phonology must be sufficiently similar to the data structures in the directly connecting modules to allow for the relevant information to be transferred. Such interface conditions are also definitely part of Marr's program, though people seem to miss this point, perhaps because it's buried in his detailed account of the visual system, e.g. pp 317 ff on transforms between coordinate systems.

Reiss and Hale often refer to the interfaces as “transducers”, but the whole of phonology is a transducer between LTM and motor control, and between perception and LTM (“From memory to speech and back”, Halle 2003). (There also seem to be some connections between action and perception that do not include a way station at memory. We ignore those things here.) Moreover, at times the SFP proposals seem to imply transducers which would have very impressive computational abilities, and which look implausible to us as direct interfaces which can be instantiated in human neural hardware (which in perceptual systems seem largely limited to affine transformations and certain changes in topology and discretization such as are accomplished in the vision system, Palmer 1999, though see Koch 1999 for an idea of how much computation a single neuron might be able to do -- a lot).

So, first, we will adopt the proposals of Jakobson, Fant and Halle 1952 (minus the acoustic definitions, which are nevertheless very relevant for engineering applications), and Halle 1983 regarding distinctive features and their neural instantiation, but drawing the feature set from Avery & Idsardi 1999. Here is a relevant diagram from Halle 1983:






And here is A&I’s wildly speculative proposal:


 
We believe that SFP (or at least the Reiss & Hale contingent) is sort of ok with this, although they remain more open on the feature set (which for them continues to include things like [±voiced] -- by the way, also in the Halle 1983 figure above -- despite Halle & Stevens 1971, Iverson & Salmons 1995 et seq).

We understand the perception side to involve neural assemblies including spectro-temporal receptive fields (STRFs, Mesgarani, David, Fritz & Shamma 2008, Mesgarani, Cheung, Johnson & Chang 2014) and a coupled dual-time-window temporal analysis (Poeppel 2003, Giraud & Poeppel 2012). On the action side, we’ll go with Bouchard, Mesgarani, Johnson & Chang 2013 and another off-the-shelf component, Guenther 2014. Much less is known about the relevant memory systems, but we’ll take stuff like Hasselmo 2012 and Murray, Wise & Graham 2017 as some starting points.

Most importantly in our opinion (we've been telegraphing this point in previous posts) we need a reasonable understanding of what “precedes” is, and we think precedes needs to be front and center in the theory. We feel that it is very unfortunate that phonological practice has often favored implicit depictions of “precedes” (as horizontal position in a diagram) rather than being explicit about it (but we’ve rehearsed these arguments before, to not much effect). We take “precedes” to be a temporal relation at the action and perception interfaces, providing the basis in the phonology for notions such as “before” and “after” and “at the same time, more or less”. (And for speech “more of less” appears to be about 50ms, Saberi & Perrott 1999, again see Poeppel & Giraud & Ghitza &co. It's not impossible for the auditory system to detect changes that are faster than this, but such rapid transitions will be encoded in phonology as features rather than as separate events.) We feel that it bears repeating here that the precedes relation in phonology is interfacing to the data structures and relations for time that are available in motor control and auditory perception and memory, which themselves are not going to capture perfectly the physical nature of time (whatever that is). That is, the acuity of precedes will be limited by things such as the fact that there is an auditory threshold for the detection of order between two physical events. (A point that Charles brought up in the comments last week.)

There are a number of ways to construct pertinent data structures and relations, but we will choose to do this in terms of events, abstract points in time (compare Carson-Berndsen’s 1998 Time-Map Phonology in which events cover spans or intervals of linear time). This will probably seem weird at first, but we believe that it leads to a better overall model. NB: This is absolutely NOT the ONLY way to go about formalizing phonology. SFP (Bale & Reiss 2018) take quite a different approach based on set theory. We will discuss the differences in a later post or two.

Within the phonology that means that we have at least:

  1. events/elements/entities, which are points in abstract time. We will use lower case letters (e, f, g, ...) to indicate these.
  2. features, which we construct as properties of events. When we don’t care what their content is we will use upper case letters to indicate them (F, G, H, …). So Fe means that event e has feature F. We will enclose specific features in brackets, following common usage in phonology, e.g. [spread]e. Notationally, [F, G]e will mean Fe AND Ge. We will drop the event variable when it’s clear in context (think Haskell point-free notation).
  3. precedes, a 2-place relation of order over events, notated e^f (e precedes f). The exact “meaning” of this relation is a little tricky given that we are not going to put many restrictions on it. For example, following Raimy 2000, we will allow “loops in time”. We're not sure that model internal relations really have any "meaning" apart from how they function inside the system and across the interfaces, but if it helps, e^f is something like "after e you can send f next" at the motor interface and "perceived e and then perceived f next" at the perceptual interface.

(We’ll do things in this way partly because of Bromberger 1988, though I [wji] still don’t think I fully understand Sylvain’s point. Also, having (1-3) allows us to steal some of Paul Pietroski’s ideas.)


So far (1-3) give us a directed multigraph (it allows self-edges and multiple edges between nodes); and with it comes with no guarantees of connectedness yet. (And Jon Rawski would like us to point out that there's a fore-shadowing of a model-theoretic approach here. Jon, please say more in the comments if you'd like.) We suppose this thing needs a name, so let’s call it Event-Feature-Precedence (EFP) Theory (would PFE be better? that could be pronounced [p͡fɛ], as in “[p͡fɛ], that’s not much of a theory”). With (1-3) we have a feature-based version of Raimy 2000 (as opposed to its original x-tier orientation), but since we will allow events to have multiple properties (features), we can recreate Raimy diagrams, such as this one for “kitty-kitty” where the symbols are the usual shorthands for combinations of features.





And now, for some random quotes about non-linear time (add more in the comments, please!):

“This time travel crap, just fries your brain like a egg.” Looper

“There is no time. Many become one.” Arrival

“Thirty-one years ago, Dick Feynman told me about his "sum over histories" version of quantum mechanics. "The electron does anything it likes," he said. "It just goes in any direction at any speed, forward or backward in time, however it likes, and then you add up the amplitudes and it gives you the wave-function." I said to him, "You're crazy." But he wasn't." Freeman Dyson
(Note: Freeman Dyson is a physicist, not a movie).
We also take properties (features) to be brain states, as in Halle 1983. Then [spread]e means that event e has the property spread glottis. (For us being brain states doesn't preclude the properties from being other things too. We mean (1-3) in a Marrian way across implementations, algorithms and problem specifications.) This discussion will sometimes be cast as if features are single neurons. This is certainly a vast over-simplification, but it will do for present purposes. For us this means that a feature (= neuron (group)) can be “activated”. (The word “feature” seems to induce a lot of confusion, so we might call these constructs fneurons, which we will insist should be pronounced [fnɚɑ̃n] without a prothetic [ɛ].) The innervation of [spread] fneuron in event e (in conjunction with the correct state of volitional control circuits) will cause a signal to be sent to the motor control system that will ultimately innervate the descending laryngeal nerve to innervate the posterior cricoarytenoid muscle (and reciprocally de-innervate the lateral cricoarytenoid muscle). On the perception side, we assume (facts not in evidence because we’re too lazy to look through all of the STRFs in Mesgarani et al 2008) that there are auditory neurons whose STRFs calculate the intensity difference between bark bands 1 and 2 as versus bark bands 3 and 4 (probably modulated by the overall spectral tilt in bark bands 5 to 10). The greater this intensity difference the more likely a [spread] fneuron is to be innervated (activated past threshold). We doubt that there’s much effect of volitional control on the perceptual side, as auditory MMNs can be observed in comatose patients, or at least in those that eventually recover, 30/33 patients in Fischer, Morlet & Giard 2000. To Charles's point last week about phonological delusions, it's well known that large positive values of voice onset time (VOT) can signal [spread] also, without the inclusion of voice quality differences in the first few pitch periods of the vowel. So the working hypothesis is that the neurons for [spread] are connected to auditory neurons with at least two kinds of STRFs, the bark-based one mentioned above, and a neuron yielding a double-on response, see various publications by Steinschneider.

As many people are aware, graphs (and multigraphs) are usually defined over sets (or bags or multisets) of vertices and edges. We haven’t included the set stuff here. Why not? The idea, for the moment at least, is that we’re calculating in a workspace, and there isn’t significant substructure in the workspace in terms of individuated phonological forms. So the workspace universe provides the set structures, such as they are, the events, the properties of the events and the relations between events. This could well be a big mistake, as it seems to preclude asking (or answering) questions like “do these two words rhyme?” as this would involve comparing sub-structures of two different phonological representations. That is, in order to evaluate rhyme (or alliteration, or …) we would have to evaluate/find a matching relation between two sub-graphs, so we would need to be able to represent two separate graphs in the workspace and know which one was which. To do this we could add labels to keep track of multiple representations, or add the extra set structure. There are (different) mathematical consequences for either move, and it isn’t at all clear which would be preferable. So we won’t do anything for now. That is, we’re wimping out on this question.

Finally, we will say once more that there are some strong similarities between this approach and Carson-Berndsen’s work. But in Time-Map Phonology time is represented using intervals on a continuous linear timeline whereas here we have discretized time instead and we allow precedence to be non-linear.

Next time: It’s Musky