Comments

Showing posts with label On Wh Movement. Show all posts
Showing posts with label On Wh Movement. Show all posts

Monday, February 15, 2016

The wonder of subjacency

I am currently teaching our Grad syntax 2 course and, not surprisingly, it focuses on the Minimalist Program (MP). Given my predilections (and the influence of Howard Lasnik) I find that one can best appreciate MP by starting with Government-Binding Theory (GB). Jairo Nunes, Kleanthes Grohmann and me used GB as backdrop to MP in our intro book (here). But every time I teach this course I become more and more impressed with the virtues of GB. It is a pretty neat little theory, and, for my money, it still provides the best set of analytical tools in linguistics. In fact, were I charged with the task of describing a new construction or writing the G of a language I would render it in a GB idiom, minimalist technology be damned. However, this is not what I wanted to write about here. Rather, I wanted to sing the praises on one particular sub-part of GB that dealt with a topic that has largely fallen out of research favor but that stands as one of the great scientific accomplishments of Generative Grammar (GG). The topic? Islands and Subjacency. What follows is why I consider it such an achievement.

As everyone knows, Chomsky’s aim in developing the theory of Subjacency (S) was to unify Ross’s islands, the latter having been discovered and described about a decade earlier. The locus classicus of this effort is On Wh Movement (OWM) where Chomsky lays out the story in gory detail.  Here’s a question: what did the unification add to Ross’s original discussion?

One thing it added was unification. Looked at theoretically, Ross’s islands are a motely, a list of domains opaque to movement. From the get-go, it was hard to believe that this list was what FL/UG coded. There has to be some underlying method. Chomsky’s goal was to find it. I do not actually recall his theoretical discontent being widely shared across the GG community (but, in my experience GG hardly ever suffers from the mental unease that poor theory regularly generates in Chomsky). At any rate, in unifying Ross’s islands, OWM tries to explain why the islands we find are the islands we have. In fact, OWM tries to tie the existence of islands to general computational considerations thereby providing what is, in retrospect, an excellent paradigm of Minimalist thinking. Here is what OWM says:

… the island constraints can be explained in terms of general and quite reasonable computational properties of formal grammar (i.e. subjacency, a property of cyclic rules that states, in effect, that transformational rules have a restricted domain of potential application; SSC, which states that only the most prominent phrase in an embedded structure is accessible to rules relating it to phrases outside; PIC, which stipulates that clauses are islands subject to the language specific escape hatch..). If this conclusion can be sustained, it will be a significant result, since such conditions as CNPC and the independent wh-island constraint seem very curious and difficult to explain on other grounds. (p. 89; On WH Movement, my emphasis).

So the list like nature of the islands becomes comprehensible when viewed from a more general computational perspective. And this is indeed a virtue.

But, and I want to emphasize this, this unification is not, as it stands, an empirical argument in favor of S. Taken at face value, what OWM demonstrates is that it is possible to unify Ross’s islands on a more rational basis, but just unifying them does not show that this unification is empirically fecund or justified.

Happily, the unification proved to be empirically very fertile indeed. OWM provides two ways that S logic could lead to the discovery of novel data.

First, it provides a general method for discovering which kinds of dependencies should be subject to island effects. OWM has a long and interesting discussion of comparative constructions and notes that given the nature of the unification proposed, comparatives should be formed by movement. This was somewhat unconventional at the time (though the work is based on some earlier work by Richie Kayne that argued for this conclusion). In fact, the most carefully worked out theory of comparatives (due to Bresnan) treated comparatives as products of a deletion operation, rather than as products of movement. If memory serves, there was quite a bit of very vigorous debate on this topic over the next little while, including at the UCSD conference where OWM was originally presented. This debate became quite heated and gave lowly grad students like me an appreciation of the old adage: when elephants fight what gets hurt is the grass. At any rate, this was one consequence of the unification that OWM emphasizes. 
A digression: could Ross’s analysis been used as a diagnostic of movement? This is, in effect, what OWM does. It assumes that if dependency D obeys islands yet allows unbounded dependency (btw, this second conjunct is a critical yet often ignored part of S reasoning) then the dependency must be the product of movement. Could Ross’s theory be interpreted in the same way? Not really. Recall, that for Ross, what makes an island and island is not the movement (movement out of islands was fine for Ross). Rather what makes an island is chopping the resumptive pronoun that movement leaves behind. In other words, for Ross, islands restrict chopping, not movement. For Chomsky, S restricts the movement and resumption is analyzed as a non-movement dependency precisely because it does not show island effects.[1] Given this, comparatives are a very good place to empirically distinguish Ross’s theory from S-theory. Why? Because comparatives have no apparent resumptive analogues like DP movement cases do (*John is taller than Bill is it/that/such). But if there are no resumptives then there can be no chopping and so no expectation of islands. This would make deletion the natural generative operation sub-serving comparatives. Thus OWM’s argument that comparatives are actually products of movement, was an empirical argument for the S view of islands.[2]
The second empirical argument for the OWM unification came from a crop of new islands. Thus, the OWM story implied that complex DPs should be islands for extraction. This implied that we should find subject islands (which we more or less do: *What do pictures of hang in the National Gallery) but also that objects should be islands (which is far less evident: What did Bill paint pictures of). OWM spends some time trying to get out from under the problems that object extraction creates. To the degree that it succeeds, then the predicted presence of subject islands is an empirical plus for S-theory.[3]
The third empirical argument in favor of S is by far the best and, if my recollection is correct, the most wow-inducing. It’s successive cyclic A’ movement. The unification of islands predicted that unbounded movement (movement that shows no island effects) is nonetheless derivationally bounded in that it is made up of a series of small bounded steps. Ross’s theory made no such prediction, Indeed, prior to S-theory there was no reason to believe it to be true. Unbounded dependencies were considered to be perfectly reasonable operations. S-theory implies that, at least for one class of dependencies (i.e. movement), such unboundedness is an illusion. This was (and is) a hell of an implication. There are many many languages where there is little evidence suggesting that anything like this is true (English being a good example of one). But, as we soon discovered (and by ‘we’ I mean GGers), it was TRUE (insert fireworks and brass bands here).
I was a grad student in Cambridge when the empirical evidence started trickling in. Jean Yves Pollock gave versions of the deservedly famous paper he co-wrote with Richie Kayne on stylistic inversion in French. If memory serves, Esther Torrego’s equally excellent paper was floating around when I was still a Cambridge denizen. As most now know, this trickle soon became a torrential stream of results with many languages providing overt evidence for cyclic Wh movement (Irish, Chamorro a.o.) At any rate, that this implication of S-theory was apparently true (or at least had non-obvious data that could be explained by it) was stunning. This is what good science does: its theories imply something unexpected and the unexpected turns out to be the case. It was great. And this was the evidence that really sold S-theory.
Let me emphasize the important argumentative structure: Unifying islands as in OWM implies that all movement, even that which does not manifest island-like properties, is local. Thus islands imply successive cyclic C to C movement. The discovery that this prediction holds is stunning confirmation of the unification of Ross’s islands in OWM and a strong confirmation of S-theory.
So, if anyone asks you what unifying islands brought to the table, successive C to C movement (or edge of domain to edge of next higher domain) is one of the biggies. It served a bit like the Syntactic Structures analysis of affix-hopping and do-support in that it sold S-theory with its aha effect and thereby made it widely accepted.

One last virtue: S-theory served as a bridge to other parts of cogsci. For example, S-theory had very natural interpretations in the context of parsing theories (E.g. Berwick and Weinberg) and learnability theories (e.g. Culicover and Wexler). S-theory served as a grammatical bridge to, IMO, the richest interaction between GG and other parts of cognition witnessed to date. In fact, like the income of most Americans, GG has receded from this high point, which, is really too bad.

So, what did S-theory add? It unified islands, allowed for a refinement of our understanding of movement, led to the postulation of new islands, implied that long movements were made up of short steps and served as a productive bridge to other parts of cognition.

And it has one last virtue of contemporary relevance. It serves (IMO) as and excellent  (maybe even the best) example we have of what linguistic theory should aim for. It is our poster child for for GGs scientific bona fides, which makes it odd that S-theory appears not to be a central part of the grad syntax curriculum anymore. I say this on the basis of very cursory investigation, actually just one or two discussions with recently minted PhDs. For the reasons noted above, this is too bad. It is a beautiful GG discovery and deserves to be regularly trumpeted as one of GGs great achievements. So, next time you are at a party and there is a lull in the conversation, remember the wonders of S-theory.



[1] Note that given current work arguing that resumption involves movement raises interesting questions about S theory. Given my partiality to this excellent idea (Demirdache is a leading exponent of this line of thinking), I think that it is worth revisiting some of the OWM assumptions, though I will refrain from doing so here.
[2] There was even independent dialectal evidence in favor of the movement analysis of comparatives: John is taller than what Bill is. The ‘what’ sure looks like a relative pronoun. This observation was due to Kayne, if I recall correctly (way to go Richie!).
[3] Wh islands another “novel” island, one that in fact Ross argued at length did not exist. As you all know, the status of Wh islands is somewhat variable cross linguistically and even among speakers of the same language. Sprouse’s thesis shows that they more or less display island-like acceptability signatures. However, whatever their status (maybe they are semantic rather than syntactic as some have argued), theoretically, they fell under the OWM unification only if one makes additional assumptions about the structure of C (how many “escape” hatches it contains). This assumption was usefully investigated empirically by Reinhart and Comorovski. They showed that the degree of freedom that S-theory allowed for was in fact empirically consequential, thus providing an indirect argument in favor of the unification along OWM lines.

Tuesday, February 11, 2014

Plato, Darwin, P&P and variation

Alex C (in the comment section here (Feb. 1)) makes a point that I’ve encountered before that I would like to comment on. He notes that Chomsky has stopped worrying about Plato’s Problem (PP) (as has much of “theoretical” linguistics as I noted in the previous post) and suggests (maybe this is too much to attribute to him, if so, sorry Alex) that this is due to Darwin’s Problems (DP) occupying center stage at present. I don’t want to argue with this factual claim, for I believe that there’s lots of truth to it (though IMO, as readers of the last several posts have no doubt gathered, theory of any kind is largely absent from current research). What I want to observe is that (1) there is a tension between PP and DP and (2) that resolving it opens an important place for theoretical speculation. IMO, one of the more interesting facets of current theoretical work is that it proposes a way of resolving this tension in an empirically interesting way. This is what I want to talk about.

First the tension: PP is the observation that the PLD the child uses in developing its G is impoverished in various ways when one compares it to the properties of Gs that children attain. PP, then, is another name for the Poverty of Stimulus Problem (POS).  Generative Grammarians have proposed to “solve” this problem by packing FL with principles of UG, many of which are very language specific (LS), at least if GB is taken as a guide to the content of FL.  By LS, I mean that the principles advert to very linguisticky objects (e.g. Subjects, tensed clauses, governors, case assigners, barriers, islands, c-command, etc) and very linguisticky operations (agreement, movement, binding, case assignment, etc.).  The idea has been that making UG rich enough and endowing it with LS innate structure will allow our theories of FL to attain explanatory adequacy, i.e. to explain how, say, Gs obey islands despite the absence of good and bad data relevant to fixing them present in the PLD. 

By now, all of this is pretty standard stuff (which is not to say that everyone buys into the scheme (Alex?)), and, for the most part, I am a big fan of POS arguments of this kind and their attendant conclusions. However, even given this, the theoretical problem that PP poses has hardly been solved. What we do have (again assuming that the POS arguments are well founded (which I do believe)) is a list of (plausibly) invariant(ish) properties of Gs and an explanation for why these can emerge in Gs in the absence of the relevant data in the PLD required to fix them. Thus, why do movement rules in a given G resist extraction from islands? Because something like the Subjacency/Barriers theory is part of every Language Acquisition Device’s (LAD) FL, that’s why.

However, even given this, what we still don’t have is an adequate account of how the variant properties of Gs emerge when planted in a particular PLD environment. Why is there V to T in French but not in English? Why do we have inverse control in Tsez but not Polish? Why wh-in-situ in Chinese but multiple wh to C in Bulgarian. The answer GB provided (and so far as I can tell, the answer still) is that FL contains parameters that can be set in different ways on the basis of PLD and the various Gs we have are the result of differential parameter setting. This is the story, but we have known for quite a while that this is less a solution to the question of how Gs emerge in all their variety than it is an explanation schema for a solution. P&P models, in other words, are not so much well worked out theories than they are part of a general recipe for a theory that were we able to cook it, would produce just the kind of FL that could provide a satisfying answer to the question of how Gs can vary so much. Moreover, as many have observed (Dresher and Janet Fodor are two notable examples, see below) there are serious problems with successfully fleshing out a P&P model.

Here are two: (i) the hope that many variant properties of Gs would hinge on fixing a small number of parameters seems increasingly empirically uncertain. Cederic Boeckx and Fritz Newmeyer have been arguing this for a while, and while their claims are debated (and by very intelligent people so, at least for a non-expert like me, the dust is still too unsettled to reach firm conclusions), it seems pretty clear that the empirical merits of earlier proposed parameterizations are less obvious than we took them to be. Indeed, there appears to some skepticism about whether there are any macro-parameters (in Baker’s sense[1]) and many of the micro-parametric proposals seem to end up restating what we observe in the data: that languages can differ. What made early macro-parameter theories interesting is the idea that differences among Gs come in largish clumps. The relation between a given parameter setting and the attested surface differences was understood as one to many. If, however, it turns out that every parameter correlates with just a single difference then the value of a parametric approach becomes quite unclear, at least so far as acquisition considerations are concerned. Why? Because it implies that surface differences are just due to differing PLD, not to the different options inherent in the structure of FL. In other words, if we end up with one parameter per surface difference then variation among Gs will not be as much of a window into the structure of FL as we thought it could be.

Here’s another problem: (ii) the likely parameters are not independent. Dresher (and friends) has demonstrated this for stress systems and Fodor (and friends) has provided analogous results for syntax.  The problem with a theory where parameters are not independent is that they make it very hard to see how acquisition could be incremental. If it turns out that the value of any parameter is conditional on the value of every other parameter (or very many others) then it would seem that we are stuck with a model in which all parameters must be set at once (i.e. instantaneous learning). This is not good! To evade this problem, we need some way of imposing independence on the parameters so that they can be set piecemeal without fear of having to re-set them later on. Both Dresher and Fodor have proposed ways of solving this independence problem (both elaborate a richer learning theory for parameter values to accommodate this problem). But, I think that it is fair to say that we are still a long way from a working solution. Moreover, the solutions provided all involve greatly enriching FL in a very LS way. This is where PP runs into DP. So let’s return to the aforementioned tension between PP and DP.

One way to solve PP is to enrich FL. The problem is that the richer and more linguistically parochial FL is, the harder it becomes to understand how it might have evolved. In other words, our standard GB tack in solving PP (LS enrichment of FL) appears to make answering DP harder. Note I say ‘appears.’ There are really two problems, and they are not equally acute. Let me explain.

As noted above, we have two things that a rich FL has been used to explain; (a) invariances characteristic of all Gs and (b) the attested variation among Gs. In a P&P model, the first ‘P’ handles (a) and the second (b). I believe that we have seen glimmers of how to resolve the tension between PP’s demands on FL versus DP’s as regards the principles part of P&P. Where things have become far more obscure (and even this might be too kind) involves the second parametric P. Here’s what I mean.

As I’ve argued in the past, one important minimalist project has been to do for the principles of GB what Chomsky did for islands and movement via the theory of subjacency in On Wh Movement (OWM). What Chomsky did in this paper is theoretically unify the disparate island effects by unifying all non-local (A’) dependency constructions by proposing that they have a common movement core (viz. move WH) subject to locality restrictions characterized by Bounding Theory (BT). This was terrifically inventive theory and aside from rationalizing/unifying Ross’s very disparate Island Effects, the combination of Move WH + BT predicted that all long movement would have to be successive cyclic (and even predicted a few more islands, e.g. subject islands and Wh-islands).[2]

But to get back to PP and DP, one way of regarding MP work over the last 20 years is as an attempt to do for GB modules what Chomsky did for Ross’s Islands. I’ve suggested this many times before but what I want to emphasize here is that this MP project is perfectly in harmony with the PP observation that we want to explain many of the invariances witnessed across Gs in terms of an innately structured FL. Here there is no real tension if this kind of unification can be realized. Why not? Because if successful we retain the GB generalizations. Just as Move WH + BT retain Ross’s generalizations, a successful unification within MP will retain GB’s (more or less) and so we can continue to tell the very same story about why Gs display the invariances attested as we did before. Thus, wrt this POS problem, there is a way to harmonize DP concerns with PP concerns. Of course, this does not mean that we will successfully manage to unify the GB modules in a Move WH + BT way, but we understand what a successful solution would look like and, IMO, we have every reason to be hopeful, though this is not the place to defend this view.

So, the principles part of P&P is, we might say, DP compatible (little joke here for the cognoscenti). The problem lies with the second P. FL on GB was understood to provide not only the principles of invariance but also to specify all the possible ways that Gs could differ. The parameters in GB were part of FL! And it is hard to see how to square this with DP given the terrific linguistic specificity of these parameters. The MP conceit has been to try and understand what Gs do in terms of one (perhaps)[3] linguistically specific operation (Merge) interacting with many general cognitive/computational operations/principles.  In other words, the aim has been to reduce the parochialism of the GB version of FL. The problem with the GB conception of parameters is that it is hard to see how to recast them in similarly general terms. All the parameters exploit notions that seem very very linguo-centric. This is especially true of micro parameters, but it is even true of macro ones. So, theoretically, parameters present a real problem for DP, and this is why the problems alluded to earlier have been taken by some (e.g. me) to suggest that maybe FL has little to say about G-variation. Moreover, it might explain why it is that, with DP becoming prominent, some of the interest in PP has seemed to wane. It is due to a dawning realization that maybe the structure of FL (our theory of UG) has little to say directly about grammatical variation and typology. Taken together PP and DP can usefully constrain our theories of FL, but mainly in licensing certain inferences about what kinds of invariances we will likely discover (indeed have discovered). However, when it comes to understanding variation, if parameters cannot be bleached of their LSity (and right now, this looks to me like a very rough road), it looks to me like they will never be made to fit with the leading ideas of MP, which are in turn driven by DP. 

So, Alex C was onto something important IMO. Linguists tend to believe that understanding variation is key to understanding FL. This is taken as virtually an article of faith. However, I am no longer so sure that this is a well founded presumption. DP provides us with some reasons to doubt that the range of variation reflects intrinsic properties of FL. If that is correct, then variation per se may me of little interest for those interested in liming the basic architecture of FL. Studying various Gs will, of course, remain a useful tool for in getting the details of the invariant principles and operations right. But, unlike earlier GB P&P models, there is at least an argument to be made (and one that I personally find compelling) that the range of G-variation has nothing whatsoever to do with the structure of FL and so will shed no light on two of the fundamental questions in Generative Grammar: what’s the structure of FL and why?[4]





[1] Though Baker, a really smart guy, thinks that there are so please don’t take me as endorsing the view that there aren’t any. I just don’t know. This is just my impression from linguist in the street interviews.
[2] The confirmation of this prediction was one of the great successes of generative grammar and the papers by, e.g. Kayne and Pollock, McCloskey, Chung, Torrego, and many others are still worth reading and re-reading. It is worth noting that the Move WH + BT story was largely driven by theoretical considerations, as Chomsky makes clear in OWM. The gratifying part is that the theory proved to be so empirically fecund.
[3] Note the ‘perhaps.’ If even merge is in the current parlance “third factor” then there is nothing taken to be linguistically special about FL.
[4] Note that this quite a bit of room for “learning” theory. For if the range of variation is not built into FL then why we see the variation we do must be due to how we acquire Gs given FL/UG.  The latter will still be important (indeed critical) in that any larning theory will have to incorporate the isolated invariances. However, a large part of the range of variation will fall outside the purview of FL. I discuss this somewhat in the last chapter of A theory if syntax for any of you with a prurient interest in such matters. See, in particular, the suggestion that we drop the switch analogy in favor of a more geometrical one.

Sunday, February 9, 2014

Where Chris Collins enters the fray

Chris sent me this longish response to some of what has appeared in the blog. In the hope of getting him to become a regularish participant in the ongoing discussions I here post his "Response to Norbert." I feel that he let me off lightly, actually. But this said, I think that I can still finds some points to disagree with. I will restrict these to the comments section and hand the floor over to him. Thx Chris.

*****

Response to Norbert

I read with interest Norbert’s recent post on formalization: “Formalization and Falsification in Generative Grammar”. Here I write some preliminary comments on his post.  I have not read other relevant posts in this sprawling blog, which I am only now learning how to navigate. So some of what I say may be redundant. 

For me the quote by Frege in the Begriffsschrift (pg. 6 of the book “Frege and Godel”) indicates what is important when he analogizes the “ideography” (basically first and second order predicate calculus) to a microscope: “But as soon as scientific goals demand great sharpness of resolution, the eye proves to be insufficient. The microscope, on the other hand, is perfectly suited to precisely such goals, but that is just why it is useless for all others.” Similarly, formalization in syntax is a tool that needs to be employed when needed. It not an absolute necessity and there are many ways of going about things (as I discuss below). By citing Frege, I am in no way claiming that we should aim at the same level of formalization that Frege did.

There is an important connection with the ideas of Rob Chametzky (posted by Norbert) in another place on this blog. As we have seen, Rob divides up theorizing into meta-theoretical, theoretical and analytical.  Analytical work, according to Chametzky is: “concerned with investigating the (phenomena of the) domain in question. It deploys and tests concepts and architecture developed in theoretical work, allowing for both understanding of the domain and sharpening of the theoretical concepts.” It is clear that more than 90% of all linguistics work (maybe 99%) is analytical, and that there is a paucity of true theoretical work.

A good example of analytical work would be Chomsky’s “On Wh-Movement”, which is one of the most beautiful and important papers in the field. Chomsky proposes the wh-diagnostics and relentlessly subjects a series of constructions to those diagnostics uncovering many interesting patterns and facts. The consequence that all these various constructions can be reduced to the single rule of “wh-movement” is a huge advanced, allowing one insight into UG. Ultimately, this paper lead to the Move-Alpha framework, and indirectly to Merge (the simplest and most general operation yet).
However, “On Wh-Movement” is what I would call “semi-formal”. It has semi-formal statements of various conditions and principles, and also lots of assumptions are left implicit. As a consequence it has the hallmark property of semi-formal work: there are no theorems and no proofs. Formalization is stating a theory clearly and formally enough that one can establish conclusively (i.e., with a proof) the relations between various aspects of the theory and between claims of the theory and claims of alternative theories.

Certainly, it would have been a waste of time to fully formalize “On Wh-Movement”. It would have expanded the text 10-20 fold at least, and added nothing. This is something that I think Pullum completely missed in his 1989 paper on formalization. The semi-formal nature of syntactic theory, also found in such classics as “Infinite Syntax” by Ross and “On Raising” by Postal, has led to a huge explosion of knowledge that people outside of linguistics/syntax cannot really understand (hence all the lame discussion out there on the internet and Facebook about what the real accomplishments of generative grammar have been), in part because syntacticians are not very good popularizers.
Theoretical work, according to Rob is:  “is concerned with developing and investigating primitives, derived concepts and architecture within a particular domain of inquiry.” There are many good examples of this kind of work in the minimalist literature. I would say Uriagereka’s original work on multi-spell-out qualifies and so does Epstein’s work on c-command, amongst others.

My feeling is that theoretical work (in Chametzky’s sense) is the natural place for formalization in linguistic theory. The reason is that it is possible, using formal assumptions to show clearly the relationship between various concepts, assumptions, operations and principles. For example, it should be possible to show, from formal work, that things like the NTC and Extension condition should really be thought of as theorems proved on the basis of assumptions about UG.  Since NTC and Extension condition are theorems, they can actually be eliminated from UG. And from this, one can wonder if that program can be extended to the full range of what syntacticians normally think about as constraints.
In this, I agree with Norbert who states: “It can lay bare what the conceptual dependencies between our basic concepts are.” Furthermore, as my previous paragraph makes clear, this mode of reasoning is particularly important for pushing the SMT forward. How can we know, with certainty, how some concept/principle/mechanism fits into the SMT? We can formalize and see if we can prove relations between our assumptions about the SMT and the various concepts/principles/mechanisms. Using the ruthless tools of definition, proof and theorem, we can gradually whittle away at UG, until we have the bare essence. I am sure that there are many surprises in store for us. Given the fundamental, abstract and subtle nature of the elements involved, such formalization is probably a necessity, if we want to avoid falling into a muddle of unclear conclusions.

A related reason for formalization (in addition to clearly stating/proving relationships between concepts and assumptions) is that it allows one to clarify murky areas. One of the biggest such areas nowadays is whether syntactic dependencies make use of chains, multi-dominance structures or something else entirely (maybe nothing else). Chomsky’s papers, including his recent ones, make references to chains at many points. But other recent work invokes multi-dominance. What are the differences and relations between these theories and are either of them really necessary? What assumptions about UG does multi-dominance or chains entail? I am afraid that without formalization it will be impossible to answer these questions. I am investigating these questions in my seminar this semester.
These questions about syntactic dependencies interact closely with TransferPF (Spell-Out) and TransferLF, which to my knowledge, have not only not been formalized but not even stated in an explicit manner. Investigating the question of whether multi-dominance, chains or some something else entirely (perhaps nothing else) is needed to model human language syntax will require a concomitant formalization of TransferPF and TransferLF, since these are the functions that make use of the structures formed by Merge.

Minimalist syntax calls for formalization in a way that previous syntactic theories did not. First, the nature of the basic operations is simple enough (e.g., Merge) to make formalization a real possibility. The baroque and varied nature of “transformations” in the “On Wh-Movement” framework and preceding work made the prospect for a full formalization more daunting.

Second, the nature of the concepts involved in minimalism, because of their simplicity and generality (e.g., copies, occurrences), are just too fundamental and subtle and abstract to resolve by talking through them in an informal or semi-formal way. With formalization we can hope to state things in such a way to make clear conceptual and empirical properties of the various proposals, and compare and evaluate them. In fact, I have recently being doing a lot of this with my colleagues, because only recently (by helping to write Collins and Stabler 2012) have I seen what the issues are.
So, in the spirit of Frege, formalization should be a tool for ordinary working syntacticians to clarify their ideas and examine them empirically and conceptually.