Comments

Saturday, December 15, 2012

Liar, Liar, Theory on Fire (Part Two)


This a follow-up to Part One, which noted that examples (1-2) present a challenge for (D).
(1)  Lari is false.
(2)  The second numbered example in this post is false.
                        (D)  for each Human Language L, there is a theory of truth
                                that can serve as the core of an adequate theory of meaning for L
To review, it seems that (2)—a.k.a. Lari—is neither true nor false. So one might hope for a truth theory that generates (3) rather than (4); 
(3)  Legit(Lari) --> [True(‘Lari is false.’) ≡ False(Lari)]
(4)  True(‘Lari is false.’) ≡ False(Lari)
where ‘Legit(_)’ is some paradox-avoiding restriction. But given “Foster’s Problem,” (D) seems to require truth theories that don’t generate (3) or (5),
(5)  True(‘Ernie snores.’) ≡ Snores(Ernie) & [True(‘Bert yells.’) ≡ Yells(Bert)]
which is true but not meaning-specifying. So one might seek a theory with a very weak system of deduction, reflecting a hypothesized competence that lets humans apply semantic axioms—lexical and compositional—only as required to compute the semantic properties of complex expressions. But can such a theory be a theory of truth of the sort required by (D)?
The worry is not that (4) is false. If Lari is neither true nor false, and in this respect like my dog Bode, then both sides of (4) are false. But if a true theory specifies a truth condition for sentence S, then in one fine sense, S—unlike Bode—has a truth condition. So one wonders: what special property does Lari lack (or have), in contrast with linguistic entities that are allegedly true-or-false? It’s no answer to say that Lari induces paradox. But following Davidson, advocates of (D) might say that truth theories for Human Languages specify truth conditions for utterances, not expressions relativized to contexts. And channeling Strawson, they might say that while utterances of (6) are typically true-or-false, utterances of (7) need not be.
            (6)  I am hungry.                             (7) Vulcan is a rocky planet.
If a user of (7) presupposes that Vulcan exists, in order to say that it is a rocky planet, that user’s utterance of (7) might fail to be false. Falsity may well require more than grammaticality and absence of truth; the world may have to “cooperate with,” or at least not frustrate, certain communicative intentions. This familiar point extends to at least some utterances of (8) and (9),
                        (8)  I saw that.                                  (9)  He is bald. 
since attempts to demonstrate an object can fail, and a speaker can wrongly assume that someone is not a vague case with respect to ‘bald’ (or ‘hungry’). So perhaps uses of (1), along with many uses of (2), fail to meet certain conditions for being true-or-false; where these conditions may themselves be determined in part by contextual factors (see, e.g., Michael Glanzberg's work).
I am happy to say that speech acts—in particular, attempts to make truth-evaluable claims by using sentences—are governed by norms of truth that go beyond any conditions specified by theories of Human Languages. It would be nice to have an ideal language whose sentences are themselves true-or-false in suitable contexts. But if Human Languages are I-languages that generate expressions in a biologically natural way, why think that theories of meaning for such languages specify truth conditions for utterances that need not be true-or-false? If truth is downstream of linguistic meaning—in that acts of using I-language sentences are candidates for being true-or-false, subject to review—why think that good theories of meaning for Human Languages will deploy the predicate ‘True(_)’? If truth is a property of utterances, it’s hard to see how specifications of truth conditions can be derived from a specification of a constrained capacity to generate meaningful expressions. (Saying that meaning is use doesn’t make it so.)
Moreover, Davidson did not discuss quantificational examples like (10) in detail.
                        (10)  I saw something.
But a Tarski-style theory of truth is specified in terms of expressions being satisfied by (or true of) sequences that assign values to variables. Given (11), T-theorems follow trivially.
                        (11) for each sentence S: True(S) ≡ for each assignment A, Satisfies(A, S)
The trick is to show how “S-theorems,” of the form indicated in (12), can be derived.
(12) for each assignment A: Satisfies(A, ‘I saw something.’) ≡ F(A)
For example, ‘F(A)’ might replaced with ‘for some assignment A* that differs from A at most with regard to A* what assigns to the variable x, A*(speaker) saw A*(x)’; where ‘A*(...)’ stands for whatever A* assigns to ‘...’, and speaker indexes a dimension of assignments that is (a la Kaplan) related to utterance interpretation via some pragmatic constraint—e.g., that assignment A* is germane to a conversational situation s only if A*(speaker) is the speaker in s.
For today, grant that instances of (12) can be meaning-specifying, despite the technicalia. The important point here is that a theory of truth for English will need to have S-theorems like these: Satisfies(A, ‘Lari is false.’) ≡ False(Lari); Satisfies(A, ‘Lari is true.’) ≡ True(Lari); Satisfies(A, ‘Lari is not true.’) ≡ ~Satisfies(A, ‘Lari is true.’); and Satisfies(A, ‘Lari is not false.’) ≡ ~Satisfies(A, ‘Lari is false.’). Such a theory will imply that no assignment satisfies (1) or (13),
            (1)  Lari is false.                                    (13)  Lari is true.
while each assignment satisfies (14) and (15). So (16) must be rejected.
                        (14)  Lari is not true.                              (15)  Lari is not false.
            (16)  for each sentence S: False(S) ≡ for each assignment A, ~Satisfies(A, S).
This is not yet a contradiction. One can say (14-15) are true, along with (17-18).
                        (17) It is not true that Lari is false.            
                        (18) It is not true that Lari is true.
Drawing on Kleene/Kripke, one can also say that False(S) ≡ True(not-S). But then (19)
                        (19) The 19th numbered example in “Liar, Liar, Theory on Fire” is not true.
is true if it isn’t true; so it isn’t true, and hence it is true. That’s not good. If each assignment satisfies (19) if and only if (19) isn’t true, then since (19) isn’t true, each assignment satisfies (19); in which case, (19)—a.k.a. Linus—is true. So it doesn’t help to say that (1) and (13), unlike (14-15) and (17-18), fail to meet certain conditions on being true-or-false. One can try to deny (20),
                        (20)  Satisfies(A, ‘Linus is not true.’) ≡ ~Satisfies(A, ‘Linus is true.’)
or at least offer a theory, perhaps formulated in terms of a hierarchy of types, that does not imply (20). But even if some such theory avoids analogous (“revenge”) paradoxes, remember that (D)
 (D)  for each Human Language L, there is a theory of truth
                                that can serve as the core of an adequate theory of meaning for L
requires truth theories whose theorems are meaning specifying. Moreover, as Parsons and Kripke remind us, contingent facts can make apparently innocuous claims into trouble-makers. It can seem that (21) is true in a context if and only if more than half of the relevant examples are false.
 (21) Most of the examples were false.
But imagine twenty-one examples: ten true (e.g., ‘2 >1’, ten false (e.g., ‘1 > 2’), and (21).
Like most things, many Human Language sentences are not true-or-false relative to each assignment of values to variables. So why think that any Human Language sentences have this remarkable property? I can’t prove that (10) is like (1) in being unlike Tarskian sentences. Perhaps there is “something about Lari” that makes it unusually unsatisfiable, and not merely unsatisfied—and not unsatisfiable in the ways that my dog is. But that hypothesis needs defense. Prima facie, satisfaction has its place in theories of truth, not in specifying linguistic meanings. And (D) seems to be a massive simplification: useful for certain purposes, but not true.

Wednesday, December 12, 2012

Does Anyone Ever Learn Anything?


Let’s follow Jimmy and Judy from birth to about five.  At birth they say precious little. At five you can’t shut them up. What happened in these five years? Answer: they learned their native language, English say. Obvious no?  Yes. But is it right? Did Judy and Jimmy learn English?  Well, to paraphrase a recent political celebrity, it all depends on what you mean by learn and English.  It seems undeniable that Judy and Jimmy developed a capacity absent (or at least invisible[1]) at birth and this capacity can be exercised to converse with some natives, the English speaking ones, but not others, the Mandarin speaking ones. However, does this imply that they learned English? 

Linguists have long understood that labels like English, French, Swahili, Mandarin, etc. are more convenience terms than terms of art (here’s where linguists mention the Weinreich quip that a language is just a dialect with an army and a navy; real cognoscenti adding sotto voce that dialects are just idiolects with epaulettes). However, recent research suggests that we have been far too cavalier about the first half of this doublet.  We can all agree that Judy and Jimmy acquired (competence in) English, but did they learn English? Recent (and, as we shall see, not so recent) research into word learning suggest that here we need to look before we leap, something, it appears, that kids do not do, at least when it comes to early word acquisition. I’ve already discussed some of the research by MSTG (here) which argues (very convincingly in my view) that the early stages of lexical acquisition do not involve the careful statistical weighing of competing alternatives but involves jumping to a conclusion mentally clutched with fierce determination and, with time, forgotten if incorrect, only to set up another ill supported jump into the lexical abyss (boy was that fun to write!). MSTG support this conclusion by considering learning situations less factitious than the contrived set-ups near and dear to the psych lab.  When kid and adults are asked to consider more natural (and hence visually busy) filmed vignettes in which the targets of lexical labeling are not clearly segregated and identified on pristine picture cards, they acquire word meanings all at once or not at all. This result is important for several reasons.

First and foremost, it provides a concrete illustration of why we should not equate acquisition with learning.  MSTG provides diagnostics of learning (multiple hypotheses, statistical weighing of these alternatives, gradual convergence on the right result) and argues that learning so understood fails to hold in more realistic contexts of lexical acquisition. Specific conclusion; in at least one demonstrable case, acquisition does not equal learning.  More general conclusion; it is an empirical question whether in any given acquisition context it is true that learning (now understood to be one mechanism among others for the acquisition of knowledge) is taking place.  In other words, Jimmy and Judy certainly acquired English but whether they learned it is entirely open for empirical grabs. Chomsky’s repeated suggestion that we understand language acquisition as a species of growth rather than learning makes an analogous point.  MSTG makes the case more crisply, I believe, by providing clear diagnostics of learning and showing that there are core cases of “learning” where these signature properties of learning are demonstrably absent.

Second, MSTG provide a rationale for why learning doesn’t hold in their examined cases. Commenting on this (here) I observed that the MSTG results suggest that in such busy contexts the prerequisites for “cross situational learning” do not exist and this is why an alternative acquisition strategy is employed. Following MSTG’s lead, I even suggested that learning requires the kind of structured hypothesis space provided by the contrived set ups that MSTG’s more realistic vignettes argue against. This constituted a kind of compromise position; learning applies where acquirers have well structured hypothesis spaces and leaping to conclusions holds where this fails to hold.[2]  However, (and learn from this you soft-hearted open-minded intellectual compromisers out there) once again Norbert’s natural generosity of spirit and desire for intellectual group-hug kumbaya moments led him astray. It seems that even this compromise position concedes too much to learning mechanisms. In a companion paper, Trueswell, Medina, Hafri and Gleitman (TMHG) extend the MSTG results to include the more stylized learning environments in which options are clearly marked and lexical targets (aka referents) are crisply identified.[3]  Even in the artificial setting of the experimental psych lab kids and adults do the darndest things!

More specifically, like MSTG, TMGH identify the quiddity of “cross situational learning” with the following mechanism:

…keeping track of multiple hypotheses about a word’s meaning across successive learning instances, and gradually converg[ing] on the correct meaning via an intersective statistical process (128).

What they demonstrate is that even in simple stylized psych lab contexts when acquirers “are placed in the initial novice state of identifying word-to-referent mappings across learning instances, evidence for such a multiple hypothesis tracking procedure is strikingly absent (128),” and learners don’t “track the cross-trial statistics in order to build the correct mappings between words and referents (129).”  What acquirers do is “make a single conjecture upon hearing the word used in context and carry that conjecture forward to be evaluated for consistency with the next observed context (129).”  If confirmed, the guess is retained, if disconfirmed, acquirers guess again as if de novo.

TMGH show this (same with MSTG) by considering the dynamics of knowledge fixation: “how learning accuracy unfolded across learning instances (130).” This is very interesting stuff and I strongly recommend that you look at the details. However, just as interesting is the little bit of history that TMGH review. TMGH acknowledge tapping into a long history of criticism of this traditional gradual/comparative conception of learning.  In the late 1950s and early 1960s first Irvin Rock (1957) and then William Estes (1960) used very analogous kinds of arguments to demonstrate that verbal learning was not learning at all but was one-trial guessing. They did not fare so well. As Roediger and Arnold (RA) (2012) put it “the verbal learning establishment rose up to smite down these new ideas (2).”  It is instructive to read the RA paper for it shows the weakness of the counter-arguments used against Rock and Estes that nonetheless carried the day.  Just like today, learning was less a hypothesized mechanism for acquisition than a definitional truth about it.[4] 

The TMHG paper ends with an interesting paradox that I would like to briefly discuss as well. Initial word learning is slow and laborious (35-140 words at 14 months) but unbelievably rapid thereafter (12,000 words by age 6, i.e. roughly 7 words per day for a little under 5 years).  Why the change? TMHG note other work (Lila’s syntactic bootstrapping hypothesis) which proposes that “the acquisition of syntax and other linguistic knowledge by toddlers and young preschoolers during this time period provides a rich database of additional constraints that permit the learning of many additional words.” It is conceivable that with this knowledge in place, cross situational learning might finally become operative, though as TMHG correctly observe, this is decidedly an empirical question and it is “plausible that a propose-but-verify word learning procedure is at work all along the course of word learning throughout most of the life cycle (150).”  It would be interesting in the extreme, in my view were either conclusion correct.

If the latter conclusion proved true (propose and verify all the way down), then language acquisition might have nothing to do with learning in the technical sense. Of course, this is consistent with grammar acquisition being a case of learning, i.e. perhaps we don’t learn words but we do learn parameters. Maybe, but I’d be very skeptical. If even word acquisition isn’t an instance of learning then it seems to me that the burden of proof that any area of language acquisition involves learning would be pretty high.

If, on the other hand, the former option is correct, (viz. that learning only kicks in when grammar is there to buttress it) then its role in accounting for language acquisition would, in my opinion, be quite modest.  Yes, learning plays a role, but really most of the action lies with the constraints that grammars place on the process. This does not mean that we should not study the extras that learning might be adding (though remember the possibility mooted in the prior paragraph), but I doubt that these cognitive titivations will generate much excitement if they only operate in restricted hypothesis spaces. Learning is interesting when there are lots of options that need sifting, not so much when the range of possible end states is highly restricted. 

MSTG and TMHG show us that terminology matters. If we call acquisition “learning” we’ve loaded the research dice. If MSTG and TMHG are right (and I’d bet quite a bit that they are: any suckers out there?) it looks like we’ve repeatedly loaded them to come up snake eyes. It’s time we cut our losses and open our minds to the possibility that ‘language learning’ like ‘Justice Roberts,’ ‘military intelligence,’ and ‘western civilization’ is an oxymoron.



[1] I include this hedge for every week another (usually French speaking) psychologist shows that the youngest kids seem to have the most prodigious knowledge. It seems that we have nothing to teach those little know-it-alls. 
[2] A kind of learning as first resort, jumping to conclusions a last resort, method. C.f. TMHG p. 130.
[3] It is doubtful that the meaning of a lexical item is simply its referent for reasons that Chomsky has belabored (sadly, quite unsuccessfully) over the years. There is far more structure to lexical items than what they refer to. However, for current purposes, this only further dramatizes the inadequacy of learning as a mechanism for lexical acquisition. 
[4] TMHG also cite another paper by Gallistel and friends that I have not yet read but will try to get hold of and blog about when I do. It argues (cited in TMHG 151), that “in most subjects, in most paradigms, the transition from a low level of responding to an asymptotic level is abrupt.” Oh my. I’ll keep you posted.

Monday, December 10, 2012

Revolutionary New Ideas Do Appear Infrequently



Another Post from Bob Berwick. We are working on getting him to be able to post on his own. Stay tuned but until then enjoy. (NH)

Some of the most memorable novels spring to mind with a single sentence: “Every happy family is alike; every unhappy family is unhappy in its own way;” “Depuis longtemps, je me suis couché de bonne heure.”  When it comes to linguistics, most would agree that top honors for most memorable goes to Chomsky’s colorless green ideas sleep furiously.  Most memorable yes, but perhaps also most misunderstood. How so? Let me explain.  Colorless green means to draw a distinction between grammatical nonsense and ungrammatical nonsense, the same words reversed: furiously sleep ideas green colorless. Now, you may have read colorless green ideas so often when leafing through Syntactic Structures that your mind simply skips right past these examples on page 16, (1) for colorless and (2), for its reversal.  Or perhaps you’ve read colorless green ideas so often that it’s started making sense to you. Or perhaps you haven’t read SS at all, but merely heard it ‘bruited in the byways’ that modern day natural language statisticians have pooh-poohed the contrast, figuring out that example (1) is roughly 200,000 times more likely than (2) – a completely satisfying New Age account as to why colorless green ideas seems more OK than furiously sleep ideas. If you’re this last sort of reader, or even the first, you might safely conclude that you needn’t pay attention to such examples anymore. But, you would be wrong. Ironically, nearly 60 years ago, Chomsky set out just about the same explanation for the contrast between (1) and (2) as the one now in vogue among some statistical folks. What’s more, it’s actually better – the ancient explanation provides empirical evidence backing this distinction, including why colorless green ideas is in fact so memorable, which the more recent account somehow left behind. In short, New Age, meet Old Age: congratulations, you’ve just re-discovered what was already known a couple generations ago.  The problem, it appears, is that same plague alluded to in my last post: when it comes to digging out the past, in computational linguistics you can check your library card at the door. While lousy scholarship has long been an endemic disease amongst AI people (Roger Schank once boasted that he never read anything because it “destroyed his imagination to think of new ideas” – but don’t get me started, or I’ll explain how another current AI paramour, ‘Bayes nets’ were invented by a second-year Harvard grad student in 1918, not, as is commonly believed, in the 1980s), the pestilence has spread far and wide (cf. the Internet Veil of Ignorance).  So, Sherman, let’s set the Wayback Machine to, oh, the years 1955-1956. And the place? Somewhere between Philadelphia and Cambridge, Massachusetts. 

We’ve arrived. Blowing the dust off a long, typed manuscript that has somehow vanished down the Orwellian memory hole, labeled “The Logical Structure of Linguistic Theory” aka LSLT (Order 91920 filmed by Harvard College Library), we turn to Chapter IV-145–IV-147, examples 17′ and 18′ – our examples colorless green… and furiously sleep. It’s these pages that Chomsky drew on for his class notes, SS.  And, crucially, the original contains a missing puzzle piece that didn’t find its way into SS – a revolutionary example, as it turns out.

On page IV-146, Chomsky observes that it is a matter of empirical fact that English speakers readily distinguish (1) from (2): “Yet any speaker of English will recognize at once that (1) is an absurd English sentence while (2) is no English sentence at all, and he will consequently give the normal intonation pattern of an English sentence to (1), but not to (2)” [examples renumbered to match SS.] This bit of empirical evidence is duly noted in SS: “a speaker of English will read (1) with a normal sentence intonation, but he will read (2) with a falling intonation on each word: in fact, with just the intonation pattern given to any sequence of unrelated words. He treats each word as a separate phrase” (1957, 16).  In other words, they parse (1) into its proper constituent phrases, with colorless green as the Subject; sleep furiously as the Verb Phrase, and so forth –  further empirical confirmation is that it’s easier to recall (1) than (2) – which is why colorless green takes honors as memorable – nobody remembers furiously sleep ideas green colorless.

But since in fact no English speaker has actually encountered either sentences (1) or (2) before, how do they know the two are different, and so give (1) normal intonation, assigning it normal syntactic structure, while pronouncing (2) as if it had no syntactic structure at all?   Clearly, English speakers must be using some other information than a literal count of occurrences in order to infer that colorless green… is OK, but not furiously sleeps...  Like what? Chomsky offers the following obvious solution – the puzzle piece that’s not in SS: “This distinction can be made by demonstrating that (1) is an instance of the sentence form Adjective-Adjective-Noun-Verb-Adverb, which is grammatical by virtue of such sentences as revolutionary new ideas appear infrequently that might well occur in normal English” (1955, IV-146; 1975:146). So let’s get this straight: When the observed frequency of a particular word string is zero, Chomsky proposes that people side-step the problem by using aggregated word classes rather than literal word frequencies, so that colorless falls together with revolutionary, green with new, and so forth. People then assign an aggregated word-class based phrase structure  to (1), so colorless green ideas’s effective probability is no longer zero, but something parasitic on revolutionary new ideas.  (In the Appendix to Chapter IV, 1955, Chomsky even offers an information-theoretic clustering algorithm to automatically construct such categories, with a worked-example, work done jointly with Peter Elias. But we won’t go there today.)   

Turning now to one modern statistical take on the same problem, what do we discover? The same solution: aggregate words into classes, and then use class-based frequencies to replace zero count word sequences. Here’s the relevant New Age excerpt: “we may approximate the conditional probability p(x,y) of occurrence of two words x and y in a given configuration as, p(x)∑C p(y|c)p(c|x)”…. “In particular, when (x,y) [are two words] we have an aggregate bigram model (Saul & Pereira, 1997), which is useful for modeling word sequences that include unseen bigrams” (Peirera, 2000:7).  Roughly then, instead of estimating the probability that word y follows the word x based on actual word counts, we use the likelihood that word x belongs to some word class c, and then use the likelihood that word y follows word class c.   So for instance, if colorless green never occurs, we instead note that colorless is in the same word class as revolutionary  – i.e., an Adjective – and calculate the likelihood that green follows an Adjective.  In turn, if we have a zero count for the pair green ideas, then we replace that with an estimate of the likelihood Adjective-ideas…and so on down the line.  And where do these word classes come from?  As SP note, when trained on newspaper text, these aggregate classes often correspond to meaningful word classes.  For example, in SP’s Table 3, p. 84, with 32 classes, class 8 consists of the words can, could, may, should, to, will, would.  The New Age canon then continues: “Using this estimate for the probability of a string and an aggregate model with C = 16 [ie., 16 different word classes - rcb] trained on newspaper text…we find that… p(colorless green..)/p(furiously sleep) ≈ 2 × 10-5” (i.e., about 200,000 times greater).  In other words, roughly speaking, the part of speech sequence Adjective-Adjective-Noun-Verb-Adverb is that much more likely than the sequence Adverb-Verb-Noun-Adjective-Adjective.
So what hath stats wrought? Two numbers, yes.  But a revolutionary new idea? Not so much. All the numbers pitch up is that I’m 200,000 times more likely to say colorless green ideas… than furiously sleep ideas... a statistical summary of my external behavior. But that’s it, and it’s not nearly enough.  This doesn’t really explain why people bore full-steam ahead on (1) and assign it right-as-rain syntactic structure, pronounced just like revolutionary new ideas and just as memorable, with (2) left hanging as a limp list of words.   The likelihood gap doesn’t – can’t – match the grammaticality gap. As a previous post put it, that’s just not the game we’re playing: “the rejection of the idea that linguistic competence is just (a possibly fancy statistical) summary of behaviors should be recognized as the linguistic version of the general Rationalist endorsement of the distinction between powers/natures/capacities and their behavioral/phenomenal effects.”  UG’s not a theory about statistically driven language regularities but about capacities. Nobody doubts that stats have some role to play in the (complex but murky) way that UG and knowledge of language and god knows what else interact so that the chance of my uttering carminative fulvose aglets murate ascarpatically works out to near zero, while for David Foster Wallace, that chance jumps by leaps and bounds.  Certainly not SS, which in the course of describing (1) and (2) explicitly endorses statistical methods as a way to model human linguistic behavior – SS fn 18, p. 17.  But I don’t give an apatropaic hoot about modeling the actual words coming out of my mouth. Rather, I want to explain what underlies my capacities.

Sunday, December 9, 2012

I Before E: Liar, Liar, Theory on Fire



Let’s say that a Human Language is a spoken or signed language that human children can naturally acquire. Let a Tarskian Language be a language for which there is a finitely specifiable theory of truth. If L is a Human Language, then for each sentence S of L, a human child has the capacities required to understand S. If L is a Tarskian Language, then there is a (true) theory such that for each sentence S of L, there is a corresponding theorem of the following form: True(S) ≡ P; where ‘≡’ is the material biconditional. Let Kaplanian Languages be those that are fundamentally like Tarskian Languages, while allowing for some context sensitivity of the sort illustrated with (1).
(1)  I wrote this.
If L is Kaplanian, then for each sentence S of L, there is a corresponding theorem of the following form: for each assignment A of values to context-sensitive elements of S,
True(S, A) ≡ F(A); where A might, for example, be such that the thing it assigns to ‘I’ wrote the thing that A assigns to ‘this’ (or a corresponding deictic index). We can also speak of Davidsonian Languages, for which there are truth theories whose theorems concern certain spatiotemporally located utterances of sentences. But at least for now, let’s not worry about any differences between Davidsonian and Kaplanian languages.
Here, I want to focus on examples like (2) and their bearing on thesis (D).
                        (2)  The second numbered example in this post is false.
                        (D)  for each Human Language L, there is a theory of truth
                             that can serve as the core of an adequate theory of meaning for L
(D) goes beyond the conjecture that Human Languages are Davidsonian/Kaplanian.
My concern is that even if one can (pace Tarski) specify theories of truth for Human Languages, a theory that consistently assigns truth conditions to examples like (2) will be too sophisticated to serve as the core of theory of meaning for a Human Language.
For simplicity, Let ‘Lari’ be a name for the second numbered example in this post (leaving it open whether examples are sentence-assignment pairs or utterances). If Lari is truth evaluable, then presumably, Lari is true if and only if the second numbered example in this post is false. In which case, if Lari is true, then: the second numbered example in this post is false; and since Lari is that example, Lari is false. But likewise, if Lari is false, then: since Lari is the second numbered example in this post, Lari is true if and only if Lari is false; hence, Lari is true. So prima facie, Lari is neither true nor false. By itself, that’s not paradoxical. Many things are neither true nor false: dogs, numbers, etc. But if a theory implies that Lari is true or false, that tells against the theory.
To be sure, there are ways of stipulating consistent truth theories for languages that generate analogs of (2) and are like English in other interesting respects. Clever logicians can invent clever theories that help us avoid certain mistakes in our attempts to describe the world. But it doesn’t follow that for each Human Language L, there is a theory of truth that is not too clever to be the core of a good theory of meaning for L.
As Davidson noted and Dummett stressed, the idea behind (D) is that given some assumptions about how truth is related to the use of expressions—including demonstratives, indexicals, and nondeclarative sentences—a suitably formulated theory of truth for L can specify what expressions of L mean by specifying how they determine the conditions in which more complex sentences of L are true. But such a theory must meet conditions that may not be jointly satisfiable. First, it has to be consistent, even given examples like (2). Second, as Foster noted and Davidson admitted, it seems that the theory will need to have theorems like (3) without having theorems like (4).
(3)  True(‘Ernie snores.’) ≡ Snores(Ernie)
                        (4)  True(‘Ernie snores.’) ≡ Snores(Ernie) & Precedes(Three, Seven)
Biconditionals like (3) may seem to "specify the meanings" of object language sentences, now ignoring assignment relativity for simplicity. But it is hard to see how a theory that yields biconditionals like (4) could serve as a theory of meaning for English, especially if theories of meaning are supposed to be theories of understanding. The derivability of (3) does not explain why ‘Ernie snores.’ means what it does—much less explain how speakers understand that sentence—if (4) is equally derivable.
Now if you “just” want a theory that specifies truth conditions, it does no harm if the theory generates boundlessly many instances of (5), so long as [...] is always true.
(5) True(‘Ernie snores.’) ≡ Snores(Ernie) & [...]
And of course, a theory that generates (3) may not generate (4). It depends on the axioms and background logic. But a theory that generates (3) and (6) will generate (7)
(6) True(‘Bert yells.’) ≡ Yells(Bert)
(7) True(‘Ernie snores.’) ≡ Snores(Ernie) & [True(‘Bert yells.’) ≡ Yells(Bert)]
if the theory permits replacement of ‘P’ with the conjunction of ‘P’ and a theorem. So if you want a theory of truth to serve as a theory of meaning, you need (i) a very weak background logic that doesn’t generate any “overly intellectual” theorems, or (ii) a way of identifying the meaning-specifying theorems. I see no plausible way of providing (ii). And I worry that adopting (i) is at odds with the logical sophistication—Kripke’s fixed points, Gupta-Belnap revision rules, or whatever—required to describe truth consistently.
In Knowledge of Meaning, Larson and Segal do offer an initially attractive version of (i). In effect, their system for deriving T-theorems only permits replacement of established equivalents: given P ≡ Q and Q ≡ R, P ≡ R; given x = y, Fx ≡ Fy. This system, which doesn’t license derivations of T-sentences like (4) or (7), reflects an interesting hypothesis about the human capacity to apply (lexical and compositional) semantic competence to particular expressions. But if derivability is logically blind, apart from replacement of established equivalents, then (8) will be as derivable as (3) and (6).
(8)  True(‘Lari is false.’) ≡ False(Lari)
Of course, (8) can be true if neither ‘Lari is false.’ nor Lari is true or false. Though if (8) follows from a true theory of meaning, then ‘Lari is false.’ is still truth evaluable. That’s not a contradiction. Perhaps ‘Lari is false.’ is truth evaluable—we know it to be true if and only if Lari is false—by virtue of a certain competence applying to it. But if this is the best defense of (D), I think that tells against the strategy of explaining meaning in terms of speakers knowing theories whose theorems specify how the world needs to be in order for sentences to be true. Developing this line of thought requires a second post. Stay tuned: the issues concern linguistic “knowledge” and closure under entailment.