Comments

Showing posts with label phonology. Show all posts
Showing posts with label phonology. Show all posts

Wednesday, April 11, 2018

Bale and Reiss formalism

Bill Idsardi & Eric Raimy

[Note: in this post ordered pairs and tuples will be enclosed in parentheses, (x,y), instead of with angle brackets, <x,y>. The Blogger platform tends to eat the angle brackets, which it interprets as malformed HTML. Yes, we could do it with HTML character entities, but that's painful to edit.]

Warning: this post is also not light bedtime reading.

We think that it will be instructive now to examine the Bale & Reiss (forthcoming; BR) formalism for phonology. Although their book is “just” an introductory text it is their laudable intent and very impressive achievement to be rigorous and didactic in building up their formalism. They start out with basic set theory, and they make it all very accessible to beginners, building it up piece by piece. Consequently the book is extremely clear on many matters that (all) other intro texts are vague or silent about. And we feel that the book is very successful in this regard. But (you guessed it) we have some qualms about their treatment of precedence.

In their formalism, BR have:
  1. Values, drawn from W = {+, -}
  2. Features, drawn from F = {high, low, … } -- a finite set
  3. Feature-value pairs, elements of S = W x F -- also finite
  4. Segments, which are consistent sets of feature-value pairs (p 377); see below -- also finite
  5. Forms, which are tuples of segments (p 36, pp 101-103) -- this is intended to be an infinite set
Notice that there is no mention of time or precedence here yet, so we’ll have to wait to see how that’s constructed (hint: it’s in 5, sort of).

Here’s the definition of Consistency (BR 445):

So consistency is a well-formedness condition on segments, for it goes above and beyond the set requirement by itself, as {(+,high), (-,high)} is a set of feature-value pairs. Like our discussion of EFP features a few posts ago, this is a NAND condition, +high NAND -high, it’s just a little harder to state now without a basic notion of events since it has to be stated as a condition on certain sets of features (absent a clear notion of time) rather than as a conjunction of properties of an event. That is, since segments are constructed as sets of features, the condition is stated as a condition on sets. The way it’s formulated, it quantifies over sets (A set of features …) and also over features inside those sets (no feature …) and consequently this is a second-order statement since it is quantifying over sets. However, since the set S is finite, the work done by this definition could instead be done by exhaustively listing the licit combinations (the consistent elements of the power set pow(S)) without using any quantifiers.

Without an available formal notion of time at this point, it’s a little difficult to know what to make of the segment datatype. The latent idea, so far unexpressed in the formalism, is that the segments are sets of feature-value pairs, maybe occurring “at the same time, more or less” or "overlapping in time, more or less" or something like that. But that’s not formally expressed yet, except by allusion in the name of the construct, “segment”. Presumably some statements in the transducers handle the relationships between the elements of a segment and their motor and auditory correlates. Therefore, without knowing what's in the transducers, it's hard to know if "being together in a segment" is a substantive notion or not. If there is some property such as approximate temporal overlap that's veridical with "in the same segment" then the notion would seem to qualify as substantive (at least as we understand it, i.e. veridical and useful). But Chomsky 1964's arguments regarding linearity are very powerful here, and strongly suggest that there is no obvious veridical notion in the overall mapping from UR to phonetics. So in that case, segments could be a purely formal part of the theory, with their in-the-same-set-edness not mapping to any consistent property in the motor or perceptual systems. That is, the transductions for segmenthood would be "interesting"; we think this is Veno's view at least.

But this will also depend on what the rule system does in the phonology, and therefore we shouldn’t be too hasty to think that the linearity arguments directly establish this point for SFP. One of Chomsky's linearity arguments was the comparison between “writer” and “rider”. With rules of flapping and vowel lengthening, the LTM (UR) distinction between /rayt+r/ and /rayd+r/ is mapped to a surface difference in the length of the preceding vowel [rayDr] and [ra:yDr] (these are Chomsky's transcriptions). So the derivation as a whole, as well as the LTM representation do not respect a condition of linearity with "phonetic" representations. But what about the output forms of the phonology, [rayDr] and [ra:yDr]? How do they fare with respect to linearity with the motor and auditory forms at the interface? Better, certainly, but are they "phonetic enough" to find a veridical relationship with specified aspects of the motor and perceptual systems, especially at the point of interface? This seems like a hard, and important question.

Moreover, if features are substantive, the +high NAND -high condition has at least a potential lawful relationship to the co-domain motor and perceptual conditions, i.e. mot(+high) NAND mot(-high) and also aud(+high) NAND aud(-high), where mot() and aud() are the transductions with the motor and auditory systems respectively. If these NAND statements are true motor and perceptual statements, then the consistency requirement is recapitulating motor and perceptual conditions within the model (as we -- er and wji -- think it probably should). But this isn’t entirely substance-free then. What would be a (purely) formal (and non-substantive) universal is if we could show that it is NOT the case that mot(+high) NAND mot(-high) and also not the case that aud(+high) NAND aud(-high). Then having +high NAND -high in the phonology would be a phonological truth without any motor or perceptual connection or motivation. And depending on the actual content of mot() and aud() that could perhaps be the case, but we need some actual proposals for mot() and aud() in order to evaluate that. If so, then +high NAND -high could be an example of a pure phonological delusion, of the type that Charles suggested last week.

Now to the forms. For much of their presentation, BR just call them strings without saying what that means formally. But we do find out on p 36 that they intend them to be what they call “ordered sets”, which they then tell us are tuples, see also BR chapter 18. We wish they hadn’t used the term “ordered set” because that term already has an established meaning in mathematics as a structure (S, R) where S is a set and R is a relation of order over the set, e.g. Schröder 2003 Ordered Sets. The more usual treatment would be to define strings inductively using concatenation (e.g. Harrison 1978).

OK, so what’s a tuple? Using tuples for this purpose brings up some very interesting issues. Counting up by size, there is only one 0-tuple, (). (Bale and Reiss use angle brackets, <>.) The 1-tuples are from S, notated (x), the 2-tuples are from S x S, notated (x,y)  -- think points on a plane -- the 3-tuples from S x S x S (x,y,z) -- think points in 3D space -- and so on. A problem here is that unless there is a fixed upper limit on the size of the tuples, then this is not finitely axiomatizable in first-order logic as it requires an infinite number of statements. This issue has a famous history in the case of arithmetic, Ryll-Nardzewski 1952, Mostowski 1952, Montague 1964. (I [wji] already commented on the blog about problems like this in regard to < and successor. You need transitive closure of successor to get <, and that’s not first-order finitely axiomatizable.) There are alternative definitions for tuples using nested tuples, but that doesn’t get us out of the problem here, which is ultimately one of inductive (recursive) definitions. This seemingly innocuous move trips up many, many people (see Keller 2004, Some Remarks on the Definability of Transitive Closure in First-order Logic and Datalog).

Also unhelpfully, the types for tuples are all different from each other, as (x,y) has nothing to do with (x,y,z). (Hutton 2016:26 is helpful here as are the discussions in formal semantics of things like transitive (e, (e,t)) and intransitive (e,t) verbs.)

So how do BR get precedence? The way they do this is to invoke a convention to index the components of the tuples (BR pp 101-1033); their discussion refers to them primarily as strings.
"These strings have an implied left-to-right linear order and are equivalent in structure to ordered sets written with angled brackets, as we discussed above. For example, the mental representation of the word man will be mᴹæᴹnᴹ which is equivalent to <mᴹ, æᴹ, nᴹ>." (p 101)
"We use various numeral subscripts, or indexes, not only to distinguish between the variables but also to indicate their relative position in a string. Thus, if x₁x₂x₃ = mᴹæᴹnᴹ, then x₁ = mᴹ (the first member of the string), x₂ = æᴹ (the second member of the string), and x₃ = nᴹ (the third member of the string)." (p 102)
As could be predicted, we're not keen about implied representations for precedence or order and would much prefer an explicit notation for it instead. We're also not sure what "left-to-right" means other than something about typography. Is this a statement about how phonological representations map to temporal relations in the motor or auditory system?

But in addition, there are a couple of other issues with doing things this way. First, we don’t have any numbers yet because (1-5) didn’t provide any. So we have to give ourselves an infinite ordered set (in the usual sense, e.g. Schröder 2003), presumably the natural numbers N, which, as we said, are also not first-order finitely axiomatizable. When we have the numbers we can get the definitions that BR give, once we actually do the tuple indexing. It’s obvious how to do it, but since we are being formal, then we still need to say it. So here it is:
  • For all forms (x,y), x has index 1, y has index 2
  • For all forms (x,y,z), x has index 1, y has index 2, z has index 3
  • For all forms (w,x,y,z), w has index 1, x has index 2, y has index 3, z has index 4
  • ...
But now we’ve got the whole set of natural numbers in phonology. Do we really want them in there? Now we can talk about things like the 25th segment, do we want to be able to do that? There is another way out, without using any numbers or indexes. We can instead define the precedence relations directly on the components of the tuples, as follows:
  • For all forms (x,y), x^y
  • For all forms (x,y,z), x^y and y^z
  • For all forms (w,x,y,z), w^x and x^y and y^z
  • ...
You get the picture. (By the way, we’re quantifying over tuples of sets of feature-value pairs here.) Now we don’t need any numbers, and we can’t talk about segments by their position indexes because there aren’t any. And, as you could guess by now, this isn’t finitely axiomatizable either, because there would be an infinite set of these statements. This is why programming languages like Haskell have datatypes like lists, and strings are then lists of symbols. Then, at least, we can write an inductive definition on the size of the list (which we could mimic here using nested tuples, which are just cons cells by another name). That’s still not a first-order finite characterization, but it seems pretty clear that we’re not going to get one by proceeding this way, going up through sets and tuples, ending up with an infinite set of types.

In summary, precedence isn’t a primitive in the BR treatment, instead they build it out of feature-value pairs, sets, tuples and indices. Doing it this way is not first-order finitely axiomatizable. But it is finitely axiomatizable if we do it with events, properties and precedence as primitives.

So what’s the upshot here? Do these arcane points about finitary vs infinitary logic really matter? Are first-order and finiteness too much to ask? Probably. We would be content with monadic second-order (MSO) definable theories (which are first order plus quantification over monadic properties, like (*10) from the Boring details post, even though we rejected (*10)). Why are we ok with MSO? For one, this seems consistent with theories of semantics that we like (Pietroski) and MSO over strings is one characterization of the set of regular (= finite-state) languages, making a connection to the sub-regular hierarchy. But if we can bring everything down to first-order, then so much the better.



Monday, April 9, 2018

Arbitrary, or dogs and cats

Bill Idsardi

In recent comments Veno has put one part of the SFP view very succinctly and clearly:


"In other words, phonology (as an aspect of the mind/brain) treats features (and other units of phonological representation) as arbitrary symbols. From the point of view of phonology, then, features are substance-free units. This, of course, does not mean that features are not related to phonetic substance, and such a conceptualization of features does not preclude the construction of a neurobiologically plausible interface theory (even spelled out in Marr’s terms)."

So I think we need to unpack what "arbitrary symbols" means here. To put my cards on the table, when I read "arbitrary" I think of two things: the use of arbitrary in mathematics (meaning "anything meeting the definition") and Saussure's arbitrariness of the sign. My worry is that this view tends to exclude an important middle, the existence of abstraction -- distinct, but non-arbitrary relationships between levels of representation, call it hidden substance (because it's useful and partly veridical). And I think such non-arbitrary relationships of abstraction play an important role in sensory systems, and constitute the "substance" (or "substantive relationship") between adjacent levels of representation (being veridical and useful). So what we're going to uncover in this blog post are a few instances of "hidden substance". A word of warning -- this post might not make for light bedtime reading. On the other hand, it might be very soporific.

Let's start with a simple example of a non-arbitrary system, the unary numeral system, or tally marks. In this system for the natural numbers, 0 is the empty string, 1 = "|", 2 = "||", 3 = "|||" and so on. This is a non-arbitrary system because "more is more": larger numbers are represented with larger representations (larger numbers correspond to larger data structures). And addition in this system is concatenation, which then automatically (and non-arbitrarily) preserves important properties like being associative and commutative. So this isn't purely Saussurean as the representations have hidden substance (or partial substance if you prefer). Onomatopoeia (sound-symbolism) is another kind of in-between case, and it might be helpful at least as an analog in understanding the point here, namely that there's still some substance (veridicality and usefulness) but it has been partly obscured by the mapping (but this analogy is imperfect, and I really don't want to discuss theories of sound-symbolism here). It is also obscured by the fact that many other mappings are much more arbitrary (e.g. Roman numerals). Perhaps another way to think about this idea would be to say that arbitariness can be put on a scale of how much of the system is done via lookup tables, and how much of the system is done by general laws of combination (see Gallistel and King 1999).

Sensory systems, even mechano- and chemo-transducers, tend to do a similar kind of abstraction in lawfully transmitting some relationships. But admittedly those cases are not nearly as clean as the unary numerals toy example. For example, the conversion done by the rod cells in the retina also abides by "more is more" -- within the operational limits more photons received means more activity. (At the low end the limit reaches down to a single photon, at the high end, the rods reach saturation pretty quickly, leaving the rods relatively less to do for humans in the modern built world.)

I think the intended use of arbitrary in Veno's quote is for something like "substitutable, interchangeable in the functions, operations and/or relations". This is also what I take to be the point of the dogs-cats argument, which is in Bridget's article that she mentioned in response to our first post. Here's Bridget quoting Daniel Currie Hall, channeling Alec Marantz.


"The phonological component does not need to know whether the features it is manipulating refer to gestures or to sounds, just as the syntactic component does not need to know whether the words it is manipulating refer to dogs or to cats; it only needs to know that the features define segments and classes of segments. The phonetic component does not need to be told whether the features refer to gestures or to sounds, because it is itself the mechanism by which the features are converted into both gestures and sounds. So it does not matter whether a feature at the interface is called [peripheral], [grave], or [low F2], because the phonological component cannot differentiate among these alternatives, and the phonetic component will realize any one of them as all three." (p 206)

I have a couple of comments about this argument. First, I think that  the appropriate comparison for phonology is semantics, not syntax, because it is semantics that connects to the CI interface. That is, the question is whether the difference between dog() and cat() is semantically substantive, not if they are syntactically distinct. But the argument is presented in various places ranging across both semantics and syntax. And to be clear, I think this observation about cats and dogs is correct about both syntax and semantics. That is, interchanging cat() for dog() doesn't affect the nature of the syntactic or semantic computations, though it might eventually end up in different truth values for particular instances (e.g. dog(laika) vs cat(laika)). (And this is not an endorsement on my part of truth-value-oriented semantics, see Pietroski, but it will do for this discussion.)

The issue I have with the argument is simply that the range of semantic examples (cats vs dogs) is too narrow to reveal the hidden semantic substance. These cases both have the same type, (e,t). The important question is not about dogs and cats but about dogs and some item with a different type, like all which has type ((e,t), ((e,t),t)). Is the difference between dogs and all important within the semantic computation? The answer would seem to be yes -- that's the whole point of having a type system. And consequently, the difference between dogs and all is again a kind of hidden substance, only now between the CI notions DOG and ALL that map to dogs and all. That is, the difference between DOG and ALL in the CI system maps to a type difference in the semantics between (e,t) and ((e,t), ((e,t),t)). Consequently we should restrict the notion of free substitution to substitution between items of the same type. For me, this interchangeability notion is not a question of being "substance-free" (although one might say that the calculations are substance-agnostic, adapting Thomas Graf's use of "content-agnostic" from his reply to Norbert's post), rather the idea seems more closely related to the notion of systematicity (Fodor and Pylyshyn 1988, Aizawa 2003).

So let's try to translate the free-substitution-within-types idea into the proposed EFP structure to try to find cases of hidden substance. The resulting claim, which we largely agree with, would say that all events are freely interchangeable with other events without affecting the nature of the computation, all features are interchangeable with other features, and all precedence relations are interchangeable with other precedence relations. Not, importantly, that there are no differences between events and features and precedence relations, which all have different type signatures (in a "truth-value" phonology they would be (e), (e,t) and (e,(e,t)) respectively). So, we'll say a little boldly, type differences indicate hidden substantive differences. But having hidden substance isn't being "substance free" though it does make it harder to spot.

Free substitutability is certainly a situation "devoutly to be wished", even when restricted to items of the same type, but I'm afraid we will still fall a little short of this ideal. Why? Because there are the dreadful "special cases", meaning more examples of hidden substance. In the EFP model, the # and % events are special cases, which means that they aren't fully interchangeable with other (ordinary) events. An automatic consequence of their special status is that the precedence relations involving # and % are not fully interchangeable with other precedence relations either. And furthermore that there are no features that apply to # and % events. (I.e. [spread]e for e = # isn't a thing, though see this exchange between Lass and Morris Halle.) One way of thinking about this is that # and % are "purely formal", but that again highlights that the difference is one of substance -- either the special events don't have any (which seems not quite right), or they have weird, special properties that are preserved across the interface ("beginning/end of form").

This is similar in some respects to math cases where certain Abelian groups are extended to form fields. In those cases we need to "special case" the identity element over addition (0) to say that it does not have a multiplicative inverse (i.e. 0 has no reciprocal, or you can't divide by 0). I think in general we are so used to these special cases that we often fail to even notice them. So let me point out a very general source of special cases. In a recursive definition, the special cases will be the base cases (= stopping cases), like "end of string" or 0. Are there any additional special cases for events, features or precedence beyond the ones just noted? Unfortunately, we suspect that there are some more. Again, the minimalist program qua program is to try to keep the special cases to a minimum, not to declare them all out-of-bounds a priori.

To bring this post to another bumper-sticker conclusion, the hidden substance cases show that arbitrary ≠ systematic ≠ substance-free. We need to keep these notions separate, because they interact in interesting ways in complex, modular systems like language.

Wednesday, April 4, 2018

A modest proposal

Bill Idsardi and Eric Raimy


[Note: the following owes tremendous debts to Jeff Heinz and Paul Pietroski, but that does NOT imply their endorsement. But they can provide endorsements or disavowals in the comments if they want to. They also have the right to remain silent, because ‘Murica.]

Taking phonology to be the mind/brain model for speech, it needs to have interfaces to at least three other systems: the articulatory motor control system (action), the auditory system (perception) and the long term memory system (memory), forming a Memory-Action-Perception (MAP) loop (Poeppel & Idsardi 2012). Likewise, signed languages will have to have interfaces to action (the motor system for the hands, arms, etc.), perception (the visual system) and to memory. As discussed in the comments last week, we think therefore that sign language phonology probably also includes spatial primitives which spoken language phonology lacks. 

So the data structures inside phonology proper must be able to effectively receive and send information across those interfaces. This condition is a basic tenet of the minimalist program. Pietroski (2003: 198 Chomsky and his critics) is on-point and unmistakable :

"Indeed, a sentence is said to be a pair of instructions -- called "PF" and LF" -- for the A/P and C/I systems. If these instructions are to be usable by the extralinguistic systems, PFs and LFs may have to respect constraints that would be arbitrary from a purely linguistic perspective."

(And see Chomsky's replies in the same book for further remarks on the importance of the interface conditions, e.g. p 275.) That is, the data structures inside phonology must be sufficiently similar to the data structures in the directly connecting modules to allow for the relevant information to be transferred. Such interface conditions are also definitely part of Marr's program, though people seem to miss this point, perhaps because it's buried in his detailed account of the visual system, e.g. pp 317 ff on transforms between coordinate systems.

Reiss and Hale often refer to the interfaces as “transducers”, but the whole of phonology is a transducer between LTM and motor control, and between perception and LTM (“From memory to speech and back”, Halle 2003). (There also seem to be some connections between action and perception that do not include a way station at memory. We ignore those things here.) Moreover, at times the SFP proposals seem to imply transducers which would have very impressive computational abilities, and which look implausible to us as direct interfaces which can be instantiated in human neural hardware (which in perceptual systems seem largely limited to affine transformations and certain changes in topology and discretization such as are accomplished in the vision system, Palmer 1999, though see Koch 1999 for an idea of how much computation a single neuron might be able to do -- a lot).

So, first, we will adopt the proposals of Jakobson, Fant and Halle 1952 (minus the acoustic definitions, which are nevertheless very relevant for engineering applications), and Halle 1983 regarding distinctive features and their neural instantiation, but drawing the feature set from Avery & Idsardi 1999. Here is a relevant diagram from Halle 1983:






And here is A&I’s wildly speculative proposal:


 
We believe that SFP (or at least the Reiss & Hale contingent) is sort of ok with this, although they remain more open on the feature set (which for them continues to include things like [±voiced] -- by the way, also in the Halle 1983 figure above -- despite Halle & Stevens 1971, Iverson & Salmons 1995 et seq).

We understand the perception side to involve neural assemblies including spectro-temporal receptive fields (STRFs, Mesgarani, David, Fritz & Shamma 2008, Mesgarani, Cheung, Johnson & Chang 2014) and a coupled dual-time-window temporal analysis (Poeppel 2003, Giraud & Poeppel 2012). On the action side, we’ll go with Bouchard, Mesgarani, Johnson & Chang 2013 and another off-the-shelf component, Guenther 2014. Much less is known about the relevant memory systems, but we’ll take stuff like Hasselmo 2012 and Murray, Wise & Graham 2017 as some starting points.

Most importantly in our opinion (we've been telegraphing this point in previous posts) we need a reasonable understanding of what “precedes” is, and we think precedes needs to be front and center in the theory. We feel that it is very unfortunate that phonological practice has often favored implicit depictions of “precedes” (as horizontal position in a diagram) rather than being explicit about it (but we’ve rehearsed these arguments before, to not much effect). We take “precedes” to be a temporal relation at the action and perception interfaces, providing the basis in the phonology for notions such as “before” and “after” and “at the same time, more or less”. (And for speech “more of less” appears to be about 50ms, Saberi & Perrott 1999, again see Poeppel & Giraud & Ghitza &co. It's not impossible for the auditory system to detect changes that are faster than this, but such rapid transitions will be encoded in phonology as features rather than as separate events.) We feel that it bears repeating here that the precedes relation in phonology is interfacing to the data structures and relations for time that are available in motor control and auditory perception and memory, which themselves are not going to capture perfectly the physical nature of time (whatever that is). That is, the acuity of precedes will be limited by things such as the fact that there is an auditory threshold for the detection of order between two physical events. (A point that Charles brought up in the comments last week.)

There are a number of ways to construct pertinent data structures and relations, but we will choose to do this in terms of events, abstract points in time (compare Carson-Berndsen’s 1998 Time-Map Phonology in which events cover spans or intervals of linear time). This will probably seem weird at first, but we believe that it leads to a better overall model. NB: This is absolutely NOT the ONLY way to go about formalizing phonology. SFP (Bale & Reiss 2018) take quite a different approach based on set theory. We will discuss the differences in a later post or two.

Within the phonology that means that we have at least:

  1. events/elements/entities, which are points in abstract time. We will use lower case letters (e, f, g, ...) to indicate these.
  2. features, which we construct as properties of events. When we don’t care what their content is we will use upper case letters to indicate them (F, G, H, …). So Fe means that event e has feature F. We will enclose specific features in brackets, following common usage in phonology, e.g. [spread]e. Notationally, [F, G]e will mean Fe AND Ge. We will drop the event variable when it’s clear in context (think Haskell point-free notation).
  3. precedes, a 2-place relation of order over events, notated e^f (e precedes f). The exact “meaning” of this relation is a little tricky given that we are not going to put many restrictions on it. For example, following Raimy 2000, we will allow “loops in time”. We're not sure that model internal relations really have any "meaning" apart from how they function inside the system and across the interfaces, but if it helps, e^f is something like "after e you can send f next" at the motor interface and "perceived e and then perceived f next" at the perceptual interface.

(We’ll do things in this way partly because of Bromberger 1988, though I [wji] still don’t think I fully understand Sylvain’s point. Also, having (1-3) allows us to steal some of Paul Pietroski’s ideas.)


So far (1-3) give us a directed multigraph (it allows self-edges and multiple edges between nodes); and with it comes with no guarantees of connectedness yet. (And Jon Rawski would like us to point out that there's a fore-shadowing of a model-theoretic approach here. Jon, please say more in the comments if you'd like.) We suppose this thing needs a name, so let’s call it Event-Feature-Precedence (EFP) Theory (would PFE be better? that could be pronounced [p͡fɛ], as in “[p͡fɛ], that’s not much of a theory”). With (1-3) we have a feature-based version of Raimy 2000 (as opposed to its original x-tier orientation), but since we will allow events to have multiple properties (features), we can recreate Raimy diagrams, such as this one for “kitty-kitty” where the symbols are the usual shorthands for combinations of features.





And now, for some random quotes about non-linear time (add more in the comments, please!):

“This time travel crap, just fries your brain like a egg.” Looper

“There is no time. Many become one.” Arrival

“Thirty-one years ago, Dick Feynman told me about his "sum over histories" version of quantum mechanics. "The electron does anything it likes," he said. "It just goes in any direction at any speed, forward or backward in time, however it likes, and then you add up the amplitudes and it gives you the wave-function." I said to him, "You're crazy." But he wasn't." Freeman Dyson
(Note: Freeman Dyson is a physicist, not a movie).
We also take properties (features) to be brain states, as in Halle 1983. Then [spread]e means that event e has the property spread glottis. (For us being brain states doesn't preclude the properties from being other things too. We mean (1-3) in a Marrian way across implementations, algorithms and problem specifications.) This discussion will sometimes be cast as if features are single neurons. This is certainly a vast over-simplification, but it will do for present purposes. For us this means that a feature (= neuron (group)) can be “activated”. (The word “feature” seems to induce a lot of confusion, so we might call these constructs fneurons, which we will insist should be pronounced [fnɚɑ̃n] without a prothetic [ɛ].) The innervation of [spread] fneuron in event e (in conjunction with the correct state of volitional control circuits) will cause a signal to be sent to the motor control system that will ultimately innervate the descending laryngeal nerve to innervate the posterior cricoarytenoid muscle (and reciprocally de-innervate the lateral cricoarytenoid muscle). On the perception side, we assume (facts not in evidence because we’re too lazy to look through all of the STRFs in Mesgarani et al 2008) that there are auditory neurons whose STRFs calculate the intensity difference between bark bands 1 and 2 as versus bark bands 3 and 4 (probably modulated by the overall spectral tilt in bark bands 5 to 10). The greater this intensity difference the more likely a [spread] fneuron is to be innervated (activated past threshold). We doubt that there’s much effect of volitional control on the perceptual side, as auditory MMNs can be observed in comatose patients, or at least in those that eventually recover, 30/33 patients in Fischer, Morlet & Giard 2000. To Charles's point last week about phonological delusions, it's well known that large positive values of voice onset time (VOT) can signal [spread] also, without the inclusion of voice quality differences in the first few pitch periods of the vowel. So the working hypothesis is that the neurons for [spread] are connected to auditory neurons with at least two kinds of STRFs, the bark-based one mentioned above, and a neuron yielding a double-on response, see various publications by Steinschneider.

As many people are aware, graphs (and multigraphs) are usually defined over sets (or bags or multisets) of vertices and edges. We haven’t included the set stuff here. Why not? The idea, for the moment at least, is that we’re calculating in a workspace, and there isn’t significant substructure in the workspace in terms of individuated phonological forms. So the workspace universe provides the set structures, such as they are, the events, the properties of the events and the relations between events. This could well be a big mistake, as it seems to preclude asking (or answering) questions like “do these two words rhyme?” as this would involve comparing sub-structures of two different phonological representations. That is, in order to evaluate rhyme (or alliteration, or …) we would have to evaluate/find a matching relation between two sub-graphs, so we would need to be able to represent two separate graphs in the workspace and know which one was which. To do this we could add labels to keep track of multiple representations, or add the extra set structure. There are (different) mathematical consequences for either move, and it isn’t at all clear which would be preferable. So we won’t do anything for now. That is, we’re wimping out on this question.

Finally, we will say once more that there are some strong similarities between this approach and Carson-Berndsen’s work. But in Time-Map Phonology time is represented using intervals on a continuous linear timeline whereas here we have discretized time instead and we allow precedence to be non-linear.

Next time: It’s Musky

Thursday, March 29, 2018

Imagine no substantive possessions

Bill Idsardi & Eric Raimy
 

Let’s return now to the beginning of the exposition. Reiss 2016:1 starts out with a Lennon-Ono riff ( https://www.rollingstone.com/music/news/yoko-ono-added-as-songwriter-on-john-lennons-imagine-w488104 , thanks Karthik!):
“Imagine a theory of phonology that makes no reference to well-formedness, repair, contrast, typology, variation, language change, markedness, ‘child phonology’, faithfulness, constraints, phonotactics, articulatory or acoustic phonetics, or speech perception.”
(I wonder if you can.) Having excluded all of this stuff he wants to argue “that something remains that is worthy of the name ‘phonology’.” Unless he’s using these terms in ways we don’t understand, there would seem to be no substance left at all, as the resulting phonology can’t make reference to the motor and perceptual interfaces and any statements about precedence relations (phonotactics) are excluded. We’re also puzzled about how one can construct a formal system without employing any well-formedness conditions (axioms). And a theory without any substance is not a theory of anything. Similarly, an interface that doesn’t effectively transmit any information between two modules is not an interface but the lack of an interface.

Put in Marrian terms (Marr 1982, you knew this was in the cards when you started reading), there have to be some linking hypotheses between the computational, algorithmic and implementational levels (Marr p. 330 "the real power of the approach lies in the integration of all three levels of attack" emphasis added), and there must be reasonable interfaces which include compatibility in data structures between any connecting sub-modules contributing to the overall solution of the problem (e.g. the different visual coordinate systems, Marr 1982:317ff).

Max Papillon tells me [wji] that I’m misreading all this, and that I’m not the target audience anyway. (I do get it that I’m not considered much of a phonologist these days, the basis of a long-running joke in the Maryland department.) Perhaps then this is all a Feyerabend 1975 ish move, providing an ascetic formalist tonic to the hedonistic excess of substance, as Feyerabend 1978:127 explains in one reply to a book review (in the section called “Conversations with Illiterates”):
“I do not say that epistemology should become anarchic or that the philosophy of science should become anarchic. I say that both disciplines should receive anarchism as medicine. Epistemology is sick, it must be cured, and the medicine is anarchy. Now medicine is not something one takes all the time. One takes it for a certain period of time, and then one stops.” (emphasis in original)
Strengthening the comparison with Feyerabend, I recall Joe Pater’s comment to Mark Hale after Mark’s talk at the MIT Phonology 2000 conference, “I know what you are. You’re a philosopher!” The Feyerabend analogy is how I understood the Hale & Reiss 2000 charge of “substance abuse”: there’s too much appeal to substance, and this should be reduced (take your medicine). As a methodological maxim "Reduce Substance!" then I'm all on board. But let's not confuse ourselves into thinking that all reference to substance can be completely eliminated, for the theory has to be about something.

Fortunately, we think there are relatively concrete proposals to be made that start right where Chomsky suggests, with features and precedence. Our proposal (tune in next time) can be read as Raimy 2000 on steroids, with dollops of Avery & Idsardi 1999, Poeppel & Idsardi 2012 and Kazanina, Bowers & Idsardi 2017. (Do people take steroids with dollops of anything? Maybe with those articles as a chaser? Sorry for the mixed metaphors.)

Let’s sum up here with an attempt at our understanding of what substance and substantive should mean in the context of developing modular theories for complicated things like speech and vision. Entities or relations in the model are substantive to the degree that they do explanatory work within the model and have lawful connections across the interfaces to entities and relations in other modules. Such things are the substance of the theory. Entities or relations in the model that do not have such lawful connections are the (purely) formal or non-substantive things (we will suggest some). But this can be a hard matter to establish in any particular case, for the lawful connections will tend to be partial rather than total. The bumper sticker version of all this is “substantive = veridical and useful”.

Next time: Swifties

Wednesday, July 2, 2014

Syntax first?

A recent paper in the proceedings of the Royal Society there is a paper by Collier, Bickel, van Schaik, Manser and Townsend (CBvSMT) (here (and a little discussion here)) that tries to address the question: what came first, phonology or syntax. Chomsky has recently focused attention on this kind of question by noting that there are evolutionary arguments for divorcing the structure of language from its use in communication. Those interested in this topic and his views on this can go to the first of his lectures, where he develops this theme at length (see here and here). At any rate, the paper linked to above develops this theme from another direction.  The "syntax" discussed is very rudimentary. However, there is an interesting observation that the paper makes. It argues that there is not much phonology in the vocalization systems of other species. There are vocal patterns, but nothing like a system of sounds that serves to differentiate different meanings. If phonology is taken to be at the service of building hefty lexicons systematically, then it seems that very few vocalizers have phonology in this sense.  In fact, as they put it, the capacity to form various phonetic patterns is not sufficient to develop a phonology OR a syntax. The suggestion in the paper is that phonology only becomes worth developing once a syntax is there to support compositionally (viz. the capacity to combine smaller meaningful parts into larger ones). Once this capacity is there, the cash value of phonology and its capacity to support a large lexicon comes into play.  Indeed, CBvSMT here even suggests that phonology is akin to the way Dehaene views reading (here); a capacity to use articulatory hardware for phonological purposes, much as the visual system and auditory system can combine to give us letters and reading. Here's what the authors say:


Like songbirds [35] and some mammal species (cetaceans [60], pinnipeds [61], elephants [62], bats [63]), humans are vocal learners capable of producing a large number of different sounds. However humans are, as far as we know, the only species that use these sounds phonologi- cally to distinguish between the meanings of two sequences. This suggests that vocal learning and the capacity to produce a large number of different sounds alone are not sufficient to induce the emergence of a phonological level.  We therefore argue that the constraints leading to the use of a phonological level are more likely to be cognitive in nature rather than linked to the production capacity of a given species. Specifically, once humans developed the cognitive capacities to memorize phonological combinations and their meanings, phonology itself could become subject to cultural, as opposed to biological, evolutionary processes [23,64]. If this is the case, it might explain why phonology in the linguistic sense is so rare in the communication systems of other species. 

Truth be told, the "syntax" they point to is very rudimentary. However, the argument form is interesting.  CBvSMT notes that there are a lot of animals, many far removed from us in evo time that vocalize but only WE speak.  As Paul Pietroski has observed (and I really hope he writes this up), this suggests that vocalization is the sort of thing that is just part and parcel of animal cognition. It's the sort of thing that animals as such can do given the right selective motivations. The availability of syntax would be just that kind of motivation for it would make having a large vocabulary all that more useful. Once one acquires a generalized combinatoric trick, having lots of units to combine (and thus a way of coding such units systematically) really becomes useful. Given that mammals are congenitally able to develop vocalization (as witnessed by the fact that so many different kinds of animals have done so (from reptiles (birds) to mice, to whales to…), the adventitious emergence of syntax would make developing phonology from vocalization a real plus. So, syntax first bringing with it semantic compositionally and then phonology to really crank this capacity up.  Or, as Chomsky might put it, language isn't sound and meaning but meaning WITH sound, the second exploiting the opportunities opened up by the first.

So, take a look. Chomsky has identified two traditions regarding the "function" of language (as if it had a function!). The dominant one is that it is a tool for communication. An older tradition thinks of it as a tool for the expression of thought.  The first tradition would seem to fit well with the idea that language emerged from vocalization in some way. The second that vocalization is a secondary effect. The CBvSMT paper addresses these issues from an angle different from Chomsky's but comes to similar conclusions. Not every day that we find two different roads to Rome.