Comments

Showing posts with label Idsardi. Show all posts
Showing posts with label Idsardi. Show all posts

Wednesday, April 11, 2018

Bale and Reiss formalism

Bill Idsardi & Eric Raimy

[Note: in this post ordered pairs and tuples will be enclosed in parentheses, (x,y), instead of with angle brackets, <x,y>. The Blogger platform tends to eat the angle brackets, which it interprets as malformed HTML. Yes, we could do it with HTML character entities, but that's painful to edit.]

Warning: this post is also not light bedtime reading.

We think that it will be instructive now to examine the Bale & Reiss (forthcoming; BR) formalism for phonology. Although their book is “just” an introductory text it is their laudable intent and very impressive achievement to be rigorous and didactic in building up their formalism. They start out with basic set theory, and they make it all very accessible to beginners, building it up piece by piece. Consequently the book is extremely clear on many matters that (all) other intro texts are vague or silent about. And we feel that the book is very successful in this regard. But (you guessed it) we have some qualms about their treatment of precedence.

In their formalism, BR have:
  1. Values, drawn from W = {+, -}
  2. Features, drawn from F = {high, low, … } -- a finite set
  3. Feature-value pairs, elements of S = W x F -- also finite
  4. Segments, which are consistent sets of feature-value pairs (p 377); see below -- also finite
  5. Forms, which are tuples of segments (p 36, pp 101-103) -- this is intended to be an infinite set
Notice that there is no mention of time or precedence here yet, so we’ll have to wait to see how that’s constructed (hint: it’s in 5, sort of).

Here’s the definition of Consistency (BR 445):

So consistency is a well-formedness condition on segments, for it goes above and beyond the set requirement by itself, as {(+,high), (-,high)} is a set of feature-value pairs. Like our discussion of EFP features a few posts ago, this is a NAND condition, +high NAND -high, it’s just a little harder to state now without a basic notion of events since it has to be stated as a condition on certain sets of features (absent a clear notion of time) rather than as a conjunction of properties of an event. That is, since segments are constructed as sets of features, the condition is stated as a condition on sets. The way it’s formulated, it quantifies over sets (A set of features …) and also over features inside those sets (no feature …) and consequently this is a second-order statement since it is quantifying over sets. However, since the set S is finite, the work done by this definition could instead be done by exhaustively listing the licit combinations (the consistent elements of the power set pow(S)) without using any quantifiers.

Without an available formal notion of time at this point, it’s a little difficult to know what to make of the segment datatype. The latent idea, so far unexpressed in the formalism, is that the segments are sets of feature-value pairs, maybe occurring “at the same time, more or less” or "overlapping in time, more or less" or something like that. But that’s not formally expressed yet, except by allusion in the name of the construct, “segment”. Presumably some statements in the transducers handle the relationships between the elements of a segment and their motor and auditory correlates. Therefore, without knowing what's in the transducers, it's hard to know if "being together in a segment" is a substantive notion or not. If there is some property such as approximate temporal overlap that's veridical with "in the same segment" then the notion would seem to qualify as substantive (at least as we understand it, i.e. veridical and useful). But Chomsky 1964's arguments regarding linearity are very powerful here, and strongly suggest that there is no obvious veridical notion in the overall mapping from UR to phonetics. So in that case, segments could be a purely formal part of the theory, with their in-the-same-set-edness not mapping to any consistent property in the motor or perceptual systems. That is, the transductions for segmenthood would be "interesting"; we think this is Veno's view at least.

But this will also depend on what the rule system does in the phonology, and therefore we shouldn’t be too hasty to think that the linearity arguments directly establish this point for SFP. One of Chomsky's linearity arguments was the comparison between “writer” and “rider”. With rules of flapping and vowel lengthening, the LTM (UR) distinction between /rayt+r/ and /rayd+r/ is mapped to a surface difference in the length of the preceding vowel [rayDr] and [ra:yDr] (these are Chomsky's transcriptions). So the derivation as a whole, as well as the LTM representation do not respect a condition of linearity with "phonetic" representations. But what about the output forms of the phonology, [rayDr] and [ra:yDr]? How do they fare with respect to linearity with the motor and auditory forms at the interface? Better, certainly, but are they "phonetic enough" to find a veridical relationship with specified aspects of the motor and perceptual systems, especially at the point of interface? This seems like a hard, and important question.

Moreover, if features are substantive, the +high NAND -high condition has at least a potential lawful relationship to the co-domain motor and perceptual conditions, i.e. mot(+high) NAND mot(-high) and also aud(+high) NAND aud(-high), where mot() and aud() are the transductions with the motor and auditory systems respectively. If these NAND statements are true motor and perceptual statements, then the consistency requirement is recapitulating motor and perceptual conditions within the model (as we -- er and wji -- think it probably should). But this isn’t entirely substance-free then. What would be a (purely) formal (and non-substantive) universal is if we could show that it is NOT the case that mot(+high) NAND mot(-high) and also not the case that aud(+high) NAND aud(-high). Then having +high NAND -high in the phonology would be a phonological truth without any motor or perceptual connection or motivation. And depending on the actual content of mot() and aud() that could perhaps be the case, but we need some actual proposals for mot() and aud() in order to evaluate that. If so, then +high NAND -high could be an example of a pure phonological delusion, of the type that Charles suggested last week.

Now to the forms. For much of their presentation, BR just call them strings without saying what that means formally. But we do find out on p 36 that they intend them to be what they call “ordered sets”, which they then tell us are tuples, see also BR chapter 18. We wish they hadn’t used the term “ordered set” because that term already has an established meaning in mathematics as a structure (S, R) where S is a set and R is a relation of order over the set, e.g. Schröder 2003 Ordered Sets. The more usual treatment would be to define strings inductively using concatenation (e.g. Harrison 1978).

OK, so what’s a tuple? Using tuples for this purpose brings up some very interesting issues. Counting up by size, there is only one 0-tuple, (). (Bale and Reiss use angle brackets, <>.) The 1-tuples are from S, notated (x), the 2-tuples are from S x S, notated (x,y)  -- think points on a plane -- the 3-tuples from S x S x S (x,y,z) -- think points in 3D space -- and so on. A problem here is that unless there is a fixed upper limit on the size of the tuples, then this is not finitely axiomatizable in first-order logic as it requires an infinite number of statements. This issue has a famous history in the case of arithmetic, Ryll-Nardzewski 1952, Mostowski 1952, Montague 1964. (I [wji] already commented on the blog about problems like this in regard to < and successor. You need transitive closure of successor to get <, and that’s not first-order finitely axiomatizable.) There are alternative definitions for tuples using nested tuples, but that doesn’t get us out of the problem here, which is ultimately one of inductive (recursive) definitions. This seemingly innocuous move trips up many, many people (see Keller 2004, Some Remarks on the Definability of Transitive Closure in First-order Logic and Datalog).

Also unhelpfully, the types for tuples are all different from each other, as (x,y) has nothing to do with (x,y,z). (Hutton 2016:26 is helpful here as are the discussions in formal semantics of things like transitive (e, (e,t)) and intransitive (e,t) verbs.)

So how do BR get precedence? The way they do this is to invoke a convention to index the components of the tuples (BR pp 101-1033); their discussion refers to them primarily as strings.
"These strings have an implied left-to-right linear order and are equivalent in structure to ordered sets written with angled brackets, as we discussed above. For example, the mental representation of the word man will be mᴹæᴹnᴹ which is equivalent to <mᴹ, æᴹ, nᴹ>." (p 101)
"We use various numeral subscripts, or indexes, not only to distinguish between the variables but also to indicate their relative position in a string. Thus, if x₁x₂x₃ = mᴹæᴹnᴹ, then x₁ = mᴹ (the first member of the string), x₂ = æᴹ (the second member of the string), and x₃ = nᴹ (the third member of the string)." (p 102)
As could be predicted, we're not keen about implied representations for precedence or order and would much prefer an explicit notation for it instead. We're also not sure what "left-to-right" means other than something about typography. Is this a statement about how phonological representations map to temporal relations in the motor or auditory system?

But in addition, there are a couple of other issues with doing things this way. First, we don’t have any numbers yet because (1-5) didn’t provide any. So we have to give ourselves an infinite ordered set (in the usual sense, e.g. Schröder 2003), presumably the natural numbers N, which, as we said, are also not first-order finitely axiomatizable. When we have the numbers we can get the definitions that BR give, once we actually do the tuple indexing. It’s obvious how to do it, but since we are being formal, then we still need to say it. So here it is:
  • For all forms (x,y), x has index 1, y has index 2
  • For all forms (x,y,z), x has index 1, y has index 2, z has index 3
  • For all forms (w,x,y,z), w has index 1, x has index 2, y has index 3, z has index 4
  • ...
But now we’ve got the whole set of natural numbers in phonology. Do we really want them in there? Now we can talk about things like the 25th segment, do we want to be able to do that? There is another way out, without using any numbers or indexes. We can instead define the precedence relations directly on the components of the tuples, as follows:
  • For all forms (x,y), x^y
  • For all forms (x,y,z), x^y and y^z
  • For all forms (w,x,y,z), w^x and x^y and y^z
  • ...
You get the picture. (By the way, we’re quantifying over tuples of sets of feature-value pairs here.) Now we don’t need any numbers, and we can’t talk about segments by their position indexes because there aren’t any. And, as you could guess by now, this isn’t finitely axiomatizable either, because there would be an infinite set of these statements. This is why programming languages like Haskell have datatypes like lists, and strings are then lists of symbols. Then, at least, we can write an inductive definition on the size of the list (which we could mimic here using nested tuples, which are just cons cells by another name). That’s still not a first-order finite characterization, but it seems pretty clear that we’re not going to get one by proceeding this way, going up through sets and tuples, ending up with an infinite set of types.

In summary, precedence isn’t a primitive in the BR treatment, instead they build it out of feature-value pairs, sets, tuples and indices. Doing it this way is not first-order finitely axiomatizable. But it is finitely axiomatizable if we do it with events, properties and precedence as primitives.

So what’s the upshot here? Do these arcane points about finitary vs infinitary logic really matter? Are first-order and finiteness too much to ask? Probably. We would be content with monadic second-order (MSO) definable theories (which are first order plus quantification over monadic properties, like (*10) from the Boring details post, even though we rejected (*10)). Why are we ok with MSO? For one, this seems consistent with theories of semantics that we like (Pietroski) and MSO over strings is one characterization of the set of regular (= finite-state) languages, making a connection to the sub-regular hierarchy. But if we can bring everything down to first-order, then so much the better.



Monday, April 9, 2018

Arbitrary, or dogs and cats

Bill Idsardi

In recent comments Veno has put one part of the SFP view very succinctly and clearly:


"In other words, phonology (as an aspect of the mind/brain) treats features (and other units of phonological representation) as arbitrary symbols. From the point of view of phonology, then, features are substance-free units. This, of course, does not mean that features are not related to phonetic substance, and such a conceptualization of features does not preclude the construction of a neurobiologically plausible interface theory (even spelled out in Marr’s terms)."

So I think we need to unpack what "arbitrary symbols" means here. To put my cards on the table, when I read "arbitrary" I think of two things: the use of arbitrary in mathematics (meaning "anything meeting the definition") and Saussure's arbitrariness of the sign. My worry is that this view tends to exclude an important middle, the existence of abstraction -- distinct, but non-arbitrary relationships between levels of representation, call it hidden substance (because it's useful and partly veridical). And I think such non-arbitrary relationships of abstraction play an important role in sensory systems, and constitute the "substance" (or "substantive relationship") between adjacent levels of representation (being veridical and useful). So what we're going to uncover in this blog post are a few instances of "hidden substance". A word of warning -- this post might not make for light bedtime reading. On the other hand, it might be very soporific.

Let's start with a simple example of a non-arbitrary system, the unary numeral system, or tally marks. In this system for the natural numbers, 0 is the empty string, 1 = "|", 2 = "||", 3 = "|||" and so on. This is a non-arbitrary system because "more is more": larger numbers are represented with larger representations (larger numbers correspond to larger data structures). And addition in this system is concatenation, which then automatically (and non-arbitrarily) preserves important properties like being associative and commutative. So this isn't purely Saussurean as the representations have hidden substance (or partial substance if you prefer). Onomatopoeia (sound-symbolism) is another kind of in-between case, and it might be helpful at least as an analog in understanding the point here, namely that there's still some substance (veridicality and usefulness) but it has been partly obscured by the mapping (but this analogy is imperfect, and I really don't want to discuss theories of sound-symbolism here). It is also obscured by the fact that many other mappings are much more arbitrary (e.g. Roman numerals). Perhaps another way to think about this idea would be to say that arbitariness can be put on a scale of how much of the system is done via lookup tables, and how much of the system is done by general laws of combination (see Gallistel and King 1999).

Sensory systems, even mechano- and chemo-transducers, tend to do a similar kind of abstraction in lawfully transmitting some relationships. But admittedly those cases are not nearly as clean as the unary numerals toy example. For example, the conversion done by the rod cells in the retina also abides by "more is more" -- within the operational limits more photons received means more activity. (At the low end the limit reaches down to a single photon, at the high end, the rods reach saturation pretty quickly, leaving the rods relatively less to do for humans in the modern built world.)

I think the intended use of arbitrary in Veno's quote is for something like "substitutable, interchangeable in the functions, operations and/or relations". This is also what I take to be the point of the dogs-cats argument, which is in Bridget's article that she mentioned in response to our first post. Here's Bridget quoting Daniel Currie Hall, channeling Alec Marantz.


"The phonological component does not need to know whether the features it is manipulating refer to gestures or to sounds, just as the syntactic component does not need to know whether the words it is manipulating refer to dogs or to cats; it only needs to know that the features define segments and classes of segments. The phonetic component does not need to be told whether the features refer to gestures or to sounds, because it is itself the mechanism by which the features are converted into both gestures and sounds. So it does not matter whether a feature at the interface is called [peripheral], [grave], or [low F2], because the phonological component cannot differentiate among these alternatives, and the phonetic component will realize any one of them as all three." (p 206)

I have a couple of comments about this argument. First, I think that  the appropriate comparison for phonology is semantics, not syntax, because it is semantics that connects to the CI interface. That is, the question is whether the difference between dog() and cat() is semantically substantive, not if they are syntactically distinct. But the argument is presented in various places ranging across both semantics and syntax. And to be clear, I think this observation about cats and dogs is correct about both syntax and semantics. That is, interchanging cat() for dog() doesn't affect the nature of the syntactic or semantic computations, though it might eventually end up in different truth values for particular instances (e.g. dog(laika) vs cat(laika)). (And this is not an endorsement on my part of truth-value-oriented semantics, see Pietroski, but it will do for this discussion.)

The issue I have with the argument is simply that the range of semantic examples (cats vs dogs) is too narrow to reveal the hidden semantic substance. These cases both have the same type, (e,t). The important question is not about dogs and cats but about dogs and some item with a different type, like all which has type ((e,t), ((e,t),t)). Is the difference between dogs and all important within the semantic computation? The answer would seem to be yes -- that's the whole point of having a type system. And consequently, the difference between dogs and all is again a kind of hidden substance, only now between the CI notions DOG and ALL that map to dogs and all. That is, the difference between DOG and ALL in the CI system maps to a type difference in the semantics between (e,t) and ((e,t), ((e,t),t)). Consequently we should restrict the notion of free substitution to substitution between items of the same type. For me, this interchangeability notion is not a question of being "substance-free" (although one might say that the calculations are substance-agnostic, adapting Thomas Graf's use of "content-agnostic" from his reply to Norbert's post), rather the idea seems more closely related to the notion of systematicity (Fodor and Pylyshyn 1988, Aizawa 2003).

So let's try to translate the free-substitution-within-types idea into the proposed EFP structure to try to find cases of hidden substance. The resulting claim, which we largely agree with, would say that all events are freely interchangeable with other events without affecting the nature of the computation, all features are interchangeable with other features, and all precedence relations are interchangeable with other precedence relations. Not, importantly, that there are no differences between events and features and precedence relations, which all have different type signatures (in a "truth-value" phonology they would be (e), (e,t) and (e,(e,t)) respectively). So, we'll say a little boldly, type differences indicate hidden substantive differences. But having hidden substance isn't being "substance free" though it does make it harder to spot.

Free substitutability is certainly a situation "devoutly to be wished", even when restricted to items of the same type, but I'm afraid we will still fall a little short of this ideal. Why? Because there are the dreadful "special cases", meaning more examples of hidden substance. In the EFP model, the # and % events are special cases, which means that they aren't fully interchangeable with other (ordinary) events. An automatic consequence of their special status is that the precedence relations involving # and % are not fully interchangeable with other precedence relations either. And furthermore that there are no features that apply to # and % events. (I.e. [spread]e for e = # isn't a thing, though see this exchange between Lass and Morris Halle.) One way of thinking about this is that # and % are "purely formal", but that again highlights that the difference is one of substance -- either the special events don't have any (which seems not quite right), or they have weird, special properties that are preserved across the interface ("beginning/end of form").

This is similar in some respects to math cases where certain Abelian groups are extended to form fields. In those cases we need to "special case" the identity element over addition (0) to say that it does not have a multiplicative inverse (i.e. 0 has no reciprocal, or you can't divide by 0). I think in general we are so used to these special cases that we often fail to even notice them. So let me point out a very general source of special cases. In a recursive definition, the special cases will be the base cases (= stopping cases), like "end of string" or 0. Are there any additional special cases for events, features or precedence beyond the ones just noted? Unfortunately, we suspect that there are some more. Again, the minimalist program qua program is to try to keep the special cases to a minimum, not to declare them all out-of-bounds a priori.

To bring this post to another bumper-sticker conclusion, the hidden substance cases show that arbitrary ≠ systematic ≠ substance-free. We need to keep these notions separate, because they interact in interesting ways in complex, modular systems like language.

Wednesday, April 4, 2018

A modest proposal

Bill Idsardi and Eric Raimy


[Note: the following owes tremendous debts to Jeff Heinz and Paul Pietroski, but that does NOT imply their endorsement. But they can provide endorsements or disavowals in the comments if they want to. They also have the right to remain silent, because ‘Murica.]

Taking phonology to be the mind/brain model for speech, it needs to have interfaces to at least three other systems: the articulatory motor control system (action), the auditory system (perception) and the long term memory system (memory), forming a Memory-Action-Perception (MAP) loop (Poeppel & Idsardi 2012). Likewise, signed languages will have to have interfaces to action (the motor system for the hands, arms, etc.), perception (the visual system) and to memory. As discussed in the comments last week, we think therefore that sign language phonology probably also includes spatial primitives which spoken language phonology lacks. 

So the data structures inside phonology proper must be able to effectively receive and send information across those interfaces. This condition is a basic tenet of the minimalist program. Pietroski (2003: 198 Chomsky and his critics) is on-point and unmistakable :

"Indeed, a sentence is said to be a pair of instructions -- called "PF" and LF" -- for the A/P and C/I systems. If these instructions are to be usable by the extralinguistic systems, PFs and LFs may have to respect constraints that would be arbitrary from a purely linguistic perspective."

(And see Chomsky's replies in the same book for further remarks on the importance of the interface conditions, e.g. p 275.) That is, the data structures inside phonology must be sufficiently similar to the data structures in the directly connecting modules to allow for the relevant information to be transferred. Such interface conditions are also definitely part of Marr's program, though people seem to miss this point, perhaps because it's buried in his detailed account of the visual system, e.g. pp 317 ff on transforms between coordinate systems.

Reiss and Hale often refer to the interfaces as “transducers”, but the whole of phonology is a transducer between LTM and motor control, and between perception and LTM (“From memory to speech and back”, Halle 2003). (There also seem to be some connections between action and perception that do not include a way station at memory. We ignore those things here.) Moreover, at times the SFP proposals seem to imply transducers which would have very impressive computational abilities, and which look implausible to us as direct interfaces which can be instantiated in human neural hardware (which in perceptual systems seem largely limited to affine transformations and certain changes in topology and discretization such as are accomplished in the vision system, Palmer 1999, though see Koch 1999 for an idea of how much computation a single neuron might be able to do -- a lot).

So, first, we will adopt the proposals of Jakobson, Fant and Halle 1952 (minus the acoustic definitions, which are nevertheless very relevant for engineering applications), and Halle 1983 regarding distinctive features and their neural instantiation, but drawing the feature set from Avery & Idsardi 1999. Here is a relevant diagram from Halle 1983:






And here is A&I’s wildly speculative proposal:


 
We believe that SFP (or at least the Reiss & Hale contingent) is sort of ok with this, although they remain more open on the feature set (which for them continues to include things like [±voiced] -- by the way, also in the Halle 1983 figure above -- despite Halle & Stevens 1971, Iverson & Salmons 1995 et seq).

We understand the perception side to involve neural assemblies including spectro-temporal receptive fields (STRFs, Mesgarani, David, Fritz & Shamma 2008, Mesgarani, Cheung, Johnson & Chang 2014) and a coupled dual-time-window temporal analysis (Poeppel 2003, Giraud & Poeppel 2012). On the action side, we’ll go with Bouchard, Mesgarani, Johnson & Chang 2013 and another off-the-shelf component, Guenther 2014. Much less is known about the relevant memory systems, but we’ll take stuff like Hasselmo 2012 and Murray, Wise & Graham 2017 as some starting points.

Most importantly in our opinion (we've been telegraphing this point in previous posts) we need a reasonable understanding of what “precedes” is, and we think precedes needs to be front and center in the theory. We feel that it is very unfortunate that phonological practice has often favored implicit depictions of “precedes” (as horizontal position in a diagram) rather than being explicit about it (but we’ve rehearsed these arguments before, to not much effect). We take “precedes” to be a temporal relation at the action and perception interfaces, providing the basis in the phonology for notions such as “before” and “after” and “at the same time, more or less”. (And for speech “more of less” appears to be about 50ms, Saberi & Perrott 1999, again see Poeppel & Giraud & Ghitza &co. It's not impossible for the auditory system to detect changes that are faster than this, but such rapid transitions will be encoded in phonology as features rather than as separate events.) We feel that it bears repeating here that the precedes relation in phonology is interfacing to the data structures and relations for time that are available in motor control and auditory perception and memory, which themselves are not going to capture perfectly the physical nature of time (whatever that is). That is, the acuity of precedes will be limited by things such as the fact that there is an auditory threshold for the detection of order between two physical events. (A point that Charles brought up in the comments last week.)

There are a number of ways to construct pertinent data structures and relations, but we will choose to do this in terms of events, abstract points in time (compare Carson-Berndsen’s 1998 Time-Map Phonology in which events cover spans or intervals of linear time). This will probably seem weird at first, but we believe that it leads to a better overall model. NB: This is absolutely NOT the ONLY way to go about formalizing phonology. SFP (Bale & Reiss 2018) take quite a different approach based on set theory. We will discuss the differences in a later post or two.

Within the phonology that means that we have at least:

  1. events/elements/entities, which are points in abstract time. We will use lower case letters (e, f, g, ...) to indicate these.
  2. features, which we construct as properties of events. When we don’t care what their content is we will use upper case letters to indicate them (F, G, H, …). So Fe means that event e has feature F. We will enclose specific features in brackets, following common usage in phonology, e.g. [spread]e. Notationally, [F, G]e will mean Fe AND Ge. We will drop the event variable when it’s clear in context (think Haskell point-free notation).
  3. precedes, a 2-place relation of order over events, notated e^f (e precedes f). The exact “meaning” of this relation is a little tricky given that we are not going to put many restrictions on it. For example, following Raimy 2000, we will allow “loops in time”. We're not sure that model internal relations really have any "meaning" apart from how they function inside the system and across the interfaces, but if it helps, e^f is something like "after e you can send f next" at the motor interface and "perceived e and then perceived f next" at the perceptual interface.

(We’ll do things in this way partly because of Bromberger 1988, though I [wji] still don’t think I fully understand Sylvain’s point. Also, having (1-3) allows us to steal some of Paul Pietroski’s ideas.)


So far (1-3) give us a directed multigraph (it allows self-edges and multiple edges between nodes); and with it comes with no guarantees of connectedness yet. (And Jon Rawski would like us to point out that there's a fore-shadowing of a model-theoretic approach here. Jon, please say more in the comments if you'd like.) We suppose this thing needs a name, so let’s call it Event-Feature-Precedence (EFP) Theory (would PFE be better? that could be pronounced [p͡fɛ], as in “[p͡fɛ], that’s not much of a theory”). With (1-3) we have a feature-based version of Raimy 2000 (as opposed to its original x-tier orientation), but since we will allow events to have multiple properties (features), we can recreate Raimy diagrams, such as this one for “kitty-kitty” where the symbols are the usual shorthands for combinations of features.





And now, for some random quotes about non-linear time (add more in the comments, please!):

“This time travel crap, just fries your brain like a egg.” Looper

“There is no time. Many become one.” Arrival

“Thirty-one years ago, Dick Feynman told me about his "sum over histories" version of quantum mechanics. "The electron does anything it likes," he said. "It just goes in any direction at any speed, forward or backward in time, however it likes, and then you add up the amplitudes and it gives you the wave-function." I said to him, "You're crazy." But he wasn't." Freeman Dyson
(Note: Freeman Dyson is a physicist, not a movie).
We also take properties (features) to be brain states, as in Halle 1983. Then [spread]e means that event e has the property spread glottis. (For us being brain states doesn't preclude the properties from being other things too. We mean (1-3) in a Marrian way across implementations, algorithms and problem specifications.) This discussion will sometimes be cast as if features are single neurons. This is certainly a vast over-simplification, but it will do for present purposes. For us this means that a feature (= neuron (group)) can be “activated”. (The word “feature” seems to induce a lot of confusion, so we might call these constructs fneurons, which we will insist should be pronounced [fnɚɑ̃n] without a prothetic [ɛ].) The innervation of [spread] fneuron in event e (in conjunction with the correct state of volitional control circuits) will cause a signal to be sent to the motor control system that will ultimately innervate the descending laryngeal nerve to innervate the posterior cricoarytenoid muscle (and reciprocally de-innervate the lateral cricoarytenoid muscle). On the perception side, we assume (facts not in evidence because we’re too lazy to look through all of the STRFs in Mesgarani et al 2008) that there are auditory neurons whose STRFs calculate the intensity difference between bark bands 1 and 2 as versus bark bands 3 and 4 (probably modulated by the overall spectral tilt in bark bands 5 to 10). The greater this intensity difference the more likely a [spread] fneuron is to be innervated (activated past threshold). We doubt that there’s much effect of volitional control on the perceptual side, as auditory MMNs can be observed in comatose patients, or at least in those that eventually recover, 30/33 patients in Fischer, Morlet & Giard 2000. To Charles's point last week about phonological delusions, it's well known that large positive values of voice onset time (VOT) can signal [spread] also, without the inclusion of voice quality differences in the first few pitch periods of the vowel. So the working hypothesis is that the neurons for [spread] are connected to auditory neurons with at least two kinds of STRFs, the bark-based one mentioned above, and a neuron yielding a double-on response, see various publications by Steinschneider.

As many people are aware, graphs (and multigraphs) are usually defined over sets (or bags or multisets) of vertices and edges. We haven’t included the set stuff here. Why not? The idea, for the moment at least, is that we’re calculating in a workspace, and there isn’t significant substructure in the workspace in terms of individuated phonological forms. So the workspace universe provides the set structures, such as they are, the events, the properties of the events and the relations between events. This could well be a big mistake, as it seems to preclude asking (or answering) questions like “do these two words rhyme?” as this would involve comparing sub-structures of two different phonological representations. That is, in order to evaluate rhyme (or alliteration, or …) we would have to evaluate/find a matching relation between two sub-graphs, so we would need to be able to represent two separate graphs in the workspace and know which one was which. To do this we could add labels to keep track of multiple representations, or add the extra set structure. There are (different) mathematical consequences for either move, and it isn’t at all clear which would be preferable. So we won’t do anything for now. That is, we’re wimping out on this question.

Finally, we will say once more that there are some strong similarities between this approach and Carson-Berndsen’s work. But in Time-Map Phonology time is represented using intervals on a continuous linear timeline whereas here we have discretized time instead and we allow precedence to be non-linear.

Next time: It’s Musky

Thursday, March 29, 2018

Imagine no substantive possessions

Bill Idsardi & Eric Raimy
 

Let’s return now to the beginning of the exposition. Reiss 2016:1 starts out with a Lennon-Ono riff ( https://www.rollingstone.com/music/news/yoko-ono-added-as-songwriter-on-john-lennons-imagine-w488104 , thanks Karthik!):
“Imagine a theory of phonology that makes no reference to well-formedness, repair, contrast, typology, variation, language change, markedness, ‘child phonology’, faithfulness, constraints, phonotactics, articulatory or acoustic phonetics, or speech perception.”
(I wonder if you can.) Having excluded all of this stuff he wants to argue “that something remains that is worthy of the name ‘phonology’.” Unless he’s using these terms in ways we don’t understand, there would seem to be no substance left at all, as the resulting phonology can’t make reference to the motor and perceptual interfaces and any statements about precedence relations (phonotactics) are excluded. We’re also puzzled about how one can construct a formal system without employing any well-formedness conditions (axioms). And a theory without any substance is not a theory of anything. Similarly, an interface that doesn’t effectively transmit any information between two modules is not an interface but the lack of an interface.

Put in Marrian terms (Marr 1982, you knew this was in the cards when you started reading), there have to be some linking hypotheses between the computational, algorithmic and implementational levels (Marr p. 330 "the real power of the approach lies in the integration of all three levels of attack" emphasis added), and there must be reasonable interfaces which include compatibility in data structures between any connecting sub-modules contributing to the overall solution of the problem (e.g. the different visual coordinate systems, Marr 1982:317ff).

Max Papillon tells me [wji] that I’m misreading all this, and that I’m not the target audience anyway. (I do get it that I’m not considered much of a phonologist these days, the basis of a long-running joke in the Maryland department.) Perhaps then this is all a Feyerabend 1975 ish move, providing an ascetic formalist tonic to the hedonistic excess of substance, as Feyerabend 1978:127 explains in one reply to a book review (in the section called “Conversations with Illiterates”):
“I do not say that epistemology should become anarchic or that the philosophy of science should become anarchic. I say that both disciplines should receive anarchism as medicine. Epistemology is sick, it must be cured, and the medicine is anarchy. Now medicine is not something one takes all the time. One takes it for a certain period of time, and then one stops.” (emphasis in original)
Strengthening the comparison with Feyerabend, I recall Joe Pater’s comment to Mark Hale after Mark’s talk at the MIT Phonology 2000 conference, “I know what you are. You’re a philosopher!” The Feyerabend analogy is how I understood the Hale & Reiss 2000 charge of “substance abuse”: there’s too much appeal to substance, and this should be reduced (take your medicine). As a methodological maxim "Reduce Substance!" then I'm all on board. But let's not confuse ourselves into thinking that all reference to substance can be completely eliminated, for the theory has to be about something.

Fortunately, we think there are relatively concrete proposals to be made that start right where Chomsky suggests, with features and precedence. Our proposal (tune in next time) can be read as Raimy 2000 on steroids, with dollops of Avery & Idsardi 1999, Poeppel & Idsardi 2012 and Kazanina, Bowers & Idsardi 2017. (Do people take steroids with dollops of anything? Maybe with those articles as a chaser? Sorry for the mixed metaphors.)

Let’s sum up here with an attempt at our understanding of what substance and substantive should mean in the context of developing modular theories for complicated things like speech and vision. Entities or relations in the model are substantive to the degree that they do explanatory work within the model and have lawful connections across the interfaces to entities and relations in other modules. Such things are the substance of the theory. Entities or relations in the model that do not have such lawful connections are the (purely) formal or non-substantive things (we will suggest some). But this can be a hard matter to establish in any particular case, for the lawful connections will tend to be partial rather than total. The bumper sticker version of all this is “substantive = veridical and useful”.

Next time: Swifties

Tuesday, March 27, 2018

Beyond Epistodome

Note: This post is NOT by Norbert. It's by Bill Idsardi and Eric Raimy. This is the first in a series of posts discussing the Substance Free Phonology (SFP) program, and phonological topics more generally.

Bill Idsardi and Eric Raimy

Before the beginning

For and against method is a fascinating book, documenting the correspondence between Paul Feyerabend and Imre Lakatos in the years just before Lakatos died. Feyerabend proposed a series of exchanges with Lakatos, with Lakatos explicating his Methodology of Scientific Research Programs, and Feyerabend taking the other side, the arguments that became Against Method. We’ll try something similar here, on the Faculty of Language blog, relating to the question of substance in phonology.

Beyond Epistodome

Note: not “Epistemodome” because phonology cares not one whit about etymology Over severalteen posts we will consider the Substance Free Phonology (SFP) program outlined by Charles Reiss and Mark Hale in a number of publications, especially Hale & Reiss 2000, 2008 and Reiss 2016, 2017. We will be concentrating mainly on Reiss 2016 ( http://ling.auf.net/lingbuzz/003087/current.pdf ). Although we agree with many of their proposals, we reject almost all of the rationales they offer for them. Because that’s such an unusual combination of views, we thought that this would be a useful forum for discussion. (And we don’t think any journal would want to publish something like this anyway.) Because this is the Faculty of Language blog (FLog? FoLog? vote in the comments!), we will start with a reading. Today’s reading is from the book of LGB, chapter 1, page 10 (Chomsky 1981):
“In the general case of theory construction, the primitive basis can be selected in any number of ways, so long as the condition of definability is met, perhaps subject to conditions of simplicity of some sort. [fn 12: See Goodman (1951).] But in the case of UG, other considerations enter. The primitive basis must meet a condition of epistemological priority. That is, still assuming the idealization to instantaneous language acquisition, we want the primitives to be concepts that can plausibly be assumed to provide a preliminary, pre-linguistic analysis of a reasonable selection of presented data, that is, to provide the primary linguistic data that are mapped by the language faculty to a grammar; relaxing the idealization to permit transitional stages, similar considerations hold. [fn 13: On this matter, see Chomsky (1975, chapter 3).] It would, for example, be reasonable to suppose that such concepts as “precedes” or “is voiced” enter into the primitive basis …” (emphasis added)
So the motto here is not “substance free”, but rather “substance first, not much of that, and not much of anything else either”. Since we’re writing this during Lent (we gave up sanity for Lent), the message of privation seems appropriate. And we are sure that the minimalist ethos is clear to this blog’s readers as well. Reiss 2016:16-7 makes a different claim:
“• phonology is epistemologically prior to phonetics…
Hammarberg (1976) leads us to see that for a strict empiricist, the somewhat rounded-lipped k of coop and the somewhat spread-lipped k of keep are very different. Given their distinctness, Hammarberg make the point, obvious yet profound, that we linguists have no reason to compare these two segments unless we have a paradigm that provides us with the category k. Our phonological theory is logically prior to our phonetic description of these two segments as “kinds of k”. So our science is rationalist. As Hammarberg also points out, the same reasoning applies to the learner -- only because of a pre-existing built-in system of categories used to parse, can the learner treat the two ‘sounds’ as variants of a category: “phonology is logically and epistemologically prior to phonetics”. Phonology provides equivalence classes for phonetic discussion.” (emphasis added)
Two claims of epistemological priority enter, one claim leaves (or maybe none). The pre-existing built-in system of categories used to parse include: (1) the features (Chomsky, Reiss 2016:18), which they both agree are substantive (Chomsky: “concepts … [that] provide a preliminary pre-linguistic analysis”, Reiss 2016:26: “This work [Hale & Reiss 2003a, 2008, 1998] accepts the existence of innate substantive features”) and (2) precedence (Chomsky; Reiss is mum on this point), also substantive. (We will get to our specific proposal in post # 3.) In the case of the learner, it’s not clear if a claim of epistemological priority can be made in either direction. In our view children have both structures: they come with innate, highly specified motor, perceptual and memory architectures along with a phonology module which has interfaces to those three entities (and probably others besides, as aspects of phonological representations are available for subsequent linguistic processing, Poeppel & Idsardi 2012, and are available in at least limited ways to introspection and metalinguistic judgements, say the central systems of Fodor 1983). The goal for the child is to learn how to transfer information among these systems for the purposes of learning and using the sound structures of the languages that they encounter. We do agree with SFP that a fruitful way of approaching this question is with a system of ordered rules within the phonological component (Halle & Bromberger 1989). In terms of evolutionary (bio-linguistic) priority, it seems blindingly clear that the supporting auditory, motor and memory systems pre-date language, and the phonology module is the new kid on the block. (Whether animal call systems because they also connect memory, action and perception are homologous to phonology is an empirical matter, see Hauser 1996.) In terms of epistemological priority for scientific investigation there would seem to be a couple ways to proceed here (Hornstein & Idsardi 2014). One is to see the human system as primates + X, essentially the evolutionary view, and ask what the minimal X is that we need to add to our last common ancestor to account for modern human abilities. The answer for phonology might be “not much” (Fitch 2018). But there’s another view, more divorced from actual biology, which tries to build things up from first principles. So in this case that would mean asking what can we conclude about any system that needs to connect memory, action, and perception systems of any sort, a “Good Old Fashioned Artificial Intelligence” (GOFAI) approach (Haugeland 1985, see https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence ). This seems to be closer to what Hale & Reiss have in mind, maybe. If so, then by this general MAP definition animal call systems would qualify as phonologies. As would a lot of other activities, including reading-writing, reading-typing, rituals (Staal 1996), dancing, kung-fu fighting, etc. (Nightmares about long ago semiotics classes ensue.) But there are problems (maybe not insurmountable) in proceeding in this way. It’s not clear that there is a general theory of sensation and perception, or of action. And what there is (e.g. Fechner/Weber laws, i.e. sensory systems do logarithms) doesn’t seem particularly helpful in the present context. We think that Gallistel 2007 is particular clear on this point:
“From a computational point of view, the notion of a general purpose learning process (for example, associative learning), makes no more sense than the notion of a general purpose sensing organ—a bump in the middle of the forehead whose function is to sense things. There is no such bump, because picking up information from different kinds of stimuli—light, sound, chemical, mechanical, and so on—requires organs with structures shaped by the specific properties of the stimuli they process. The structure of an eye—including the neural circuitry in the retina and beyond—reflects in exquisite detail the laws of optics and the exigencies of extracting information about the world from reflected light. The same is true for the ear, where the exigencies of extracting information from emitted sounds dictates the many distinctive features of auditory organs. We see with eyes and hear with ears—rather than sensing through a general purpose sense organ--because sensing requires organs with modality-specific structure.” (emphasis added)
So our take on this is that we’re going to restrict the term phonology to humans for now, and so we will need to investigate the human systems for memory, action and perception in terms of their roles in human language, in order to be able to understand the interfaces. But we agree with the strategy of finding a small set of primitives (features and precedence) that we can map across the memory-action-perception (MAP) interfaces and seeing how far we can get with that inside phonology. With Fitch 2018 The phonological continuity hypothesis, though, we will consider the properties of phonology-like systems (especially auditory pattern recognition) in other animals, such as ferrets and finches to be informative about human phonology (see also Yip 2013, Samuels 2015). How much phonological difference does it make that ASL is signed-viewed instead of spoken-heard? Maybe none or maybe a lot, probably some. The idea that there would be action-perception features in both cases seems perfectly fine, though they would obviously be connecting different things (e.g. joint flexion/extension and object-centered angle in signed language and orbicularis oris activation and FM sweep in spoken languages). Does it matter that object-centered properties are further along the cortical visual processing stream (Perirhinal cortex) whereas FM sweeps are identifiable in primary auditory cortex (A1)? Can we ignore the sub-cortical differences between the visual pathway to V1 (simple) and the ascending auditory pathway to A1 (complex)? Does it matter that V1 is two dimensional (retinotopic) and so computations there have access to notions such as spatial frequency that don’t have any clear correlates in the auditory system? Do we need to add spatial relations between features to the precedence relation in our account of ASL? (This seems to be almost certainly yes.) Again, we agree that it’s a good tactic to go as far as we can with features and precedence in both cases, but we won’t be surprised if we end up explanatorily short, especially for ASL. To address a technical point, can you learn equivalence classes? Yes, you can, that’s what unsupervised learning and cluster analysis algorithms do (Hastie, Tibshirani & Freidman 2001). Those techniques aren’t free of assumptions either (No Free Lunch Theorems, Wolpert 1996), but given some reasonable starting assumptions (innate or otherwise) they do seem relevant to human speech category formation (Dillon, Dunbar & Idsardi 2013, see also Chandrasekaran, Koslov & Maddox 2014) even if we ultimately restrict this to feature selection or (de-)activation instead of feature “invention”. Next time: Just my imagination (running away with me)

Monday, November 26, 2012

Merging Birds


In the last several years I have become a really big fan of singing mice.  It seems that unbeknownst to us, these white little fur balls have been plunging from aria to aria while gorging on food pellets and simultaneously training their ever-vigilant grad student minders to react appropriately whenever they pressed a bar.  Their songs sound birdish though at a higher pitch. Now it seems that many kinds of mice sing, not only those complaining of incarceration. I was delighted and amazed (though as my daughter pointed out, we’ve known since the first Feival film that mice are great singers). 

I don’t know how extensively rodent operettas have been studied, but recently there has been a lot of research on the structure of bird song and interesting speculation about what it may tell us about the species specificity of the kind of hierarchical recursion we find in natural language (NL). Berwick, Beckers, Okanoya and Bolhuis (BBOB; hmm, kind of a stuttering version of Berwick’s first name) provide an extensive linguist friendly review of the relevant literature which I recommend to the ornithophile with interests in UG. 

BBOB’s review is especially relevant to anyone interested in the evolution of the faculty of language (FL) (ahem, I’m talking to all you minimalists out there!). They note “many striking parallels between speech and vocal production and learning in birds and humans” but also note qualitative differences “when one compares language syntax and birdsong more generally (5/1).” The value of the review, however, is not in these broad conclusions but in the detailed comparisons between phonological vs syntactic vs birdsong structure that it outlines. In particular, both birdsong and the human sound system display precedence based dependencies (1st order markov), adjacency-based dependencies, some (limited) non-adjacent dependencies, and the grouping of elements into “chunks” (“phrases,” “syllables”).  In effect, birdsongs seem restricted to linear precedence relations alone, just what Heinz and Idsardi propose suffices to represent the essentials of the human sound system. Importantly, there is no evidence that birdsong allows for the kind of hierarchical recursion that is typical of syntactic structures:

Birdsong does not admit such extended self-nested structures, even in the nightingale song chunks are not contained within other song chunks, or song packets within other song packets or contexts within contexts (5/6) (my emphasis).

Nor do they provide any evidence for unbounded dependencies, unboundedly hierarchical asymmetric “phrases,” or displacement relations (aka movement), all characteristic features of NLs.

The BBOB paper also contains an interesting comparison of songbird and human brains remarking on various possible shared vocalization homologies in human and bird brain architecture. Even FoxP2, (that ubiquitous rascal) makes a cameo appearance, with BBOB briefly reviewing the current speculations concerning how “this system may be part of a “molecular toolkit that is essential for sensory-guided motor learning” in the relevant regions of songbirds and humans (5/9).”

All in all then I found this a very useful guide to the current state of the art, especially for those with minimalist interests.

Why minimalists in particular? Because it has possible bearing on a currently active speculation regarding the species specificity and domain specificity of Merge.  Merge, recall, is the minimalist replacement for phrase structure rules (and movement). It’s the operation responsible both for unbounded hierarchical embedding and displacement.  So if birdsong displays context free patterns one source for this could be the presence of Merge as a basic operation in the songbird brain. BBOB carefully review the evidence that birdsong patterns exceed the descriptive power of finite transition networks and demand the resources of context free grammars. They conclude that there is currently “no compelling evidence” that they do (5/14). Furthermore, BBOB note that there is no evidence for displacement-like operations in birdsong, the second product of a merge-like operation. Thus, at this time, NLs alone provide clear evidence of context free and displacement structures. So, if Merge is the operation that generates such structures, there is currently no evidence that Merge has arisen in any species other than humans or in any domain other than syntax.

Why is this important for minimalists? The minimalist Genesis story goes as follows: Some “miracle” occurred in the last 100,000 years that allowed for NLs to arise in humans. Following Chomsky, let’s call this miracle “Merge.” By hypothesis, Merge is a very “simple” addition to the cognitive repertoire. Conceptually, there are (at least) two ways it might have been added: (i) Merge is a linguistically specific miracle or (ii) it is a more general cognitive one. If (ii), then we might expect Merge to have arisen before in other species and to be expressed in other cognitive domains, e.g. birdsong.  This is where BBOB’s conclusions are important for they indicate that there is currently no evidence in birdsong for the kind of structures (i.e. ones displaying unbounded nested dependencies and displacement) Merge would generate. Thus, at present, the only cognitive products of Merge we have found occur in species that have NLs, i.e. us.

Moreover, as BBOB emphasize the impact of Merge is only visible in a subpart of our linguistic products. It is a property of syntactic structures not phonological ones.  Indeed, as BBOB show, human sound systems and birdsong systems look very similar.  This suggests that Miracle Merge is quite a picky operation, exercising its powers in just a restricted part of FL (widely construed).  So not only is Merge not cognitively general, it’s not even linguistically general. Its signature properties are restricted to syntactic structures.

If this is correct, then it suggests (to me at least) that Merge is a linguistically local miracle and so proprietary to FL and so part of UG. This, I believe, comports more with Chomsky’s earlier conception of Merge, than his current one.  The former sees the capacity to build bigger and bigger hierarchically embedded structures (and movement) as resting on being able to spread “edge features” (EF) from lexical items to the complexes of lexical items that Merge forms.  So given two lexical items (LI) (each with an inherent EF), a complex inherits an EF (presumably from one of its participants) and this inherited EF is what licenses the further merging of the created complex with other EF bearing elements (LIs and earlier products of Merge). Inherited EFs then are essentially the products of labeling (full disclosure: I confess to liking this idea as I outlined/adopted a version of it here (Btw, it makes a wonderful stocking stuffer so buy early buy often!) and labeling is the miracle primarily responsible for the e(/I)mergence (like that?) of both phrase structure and displacement.

Chomsky’s more current view seems to be that labeling (and so EFs) are dispensable and that Merge alone is the source of phrase structure and movement. There is no need for EFs as Merge is defined as being able to apply to any cognitive objects at all, primitive or constructed.  In particular, both lexical items and complexes of lexical items formed by prior applications of Merge are in the domain of Merge. EFs are unnecessary and so, with a hat tip to Ockham, should be dispensed with. 

And this brings us back to birds, their songs and their brains.  It would have been a powerful piece of evidence in favor of this latter conception were a signature of merge attested in the cognitive products of some other species for it would have been evidence that the operation isn’t FL/UG peculiar.  Birdsong was a plausible place to look and it appears that it isn’t there.  BBOB’s review locates the effects of Merge exclusively to the syntax of NL.  Were Merge more domain general and less species specific we might have expected other dogs to bark (or sing more complex songs). And though absence of evidence should not be mistaken for evidence of absence, at least right now, it looks like Merge is very domain specific, something more compatible with Chomsky’s first version of Merge than his second.